README.md
14.4 KB · 275 lines · markdown Raw
1 ---
2 license: other
3 license_name: qwen-research
4 license_link: LICENSE
5 base_model: Qwen/Qwen-Image-2.1
6 base_model_relation: adapter
7 library_name: diffusers
8 pipeline_tag: text-to-image
9 tags:
10 - diffusers
11 - lora
12 - text-to-image
13 - image-to-image
14 - image-editing
15 - distillation
16 - turbo
17 - few-step
18 - qwen-image
19 - comfyui
20 - gguf
21 - int8
22 - fp8
23 - quantized
24 ---
25
26 # Qwen-Image-2.1-viggle-turbo — v0.3
27
28 **Built with Qwen.** A few-step distilled version of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1)
29 by Viggle. Text-to-image and instruction-driven editing with 1–3 reference images in **6 steps instead of 40**, with
30 **no classifier-free guidance**.
31
32 <video controls autoplay muted loop playsinline width="100%"
33 src="https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/resolve/main/assets/viggle_turbo_promo.mp4"></video>
34
35 About **5× faster** than the 40-step base model end to end, and very competitive with it in quality: on the official
36 Qwen examples the two are hard to tell apart on most prompts. The clearest gaps are small, dense text and complicated
37 edits ([Known limitations](#known-limitations)). Compare them yourself in the **Comparison** tab of the
38 [demo Space](https://huggingface.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo).
39
40 ## v0.3 (2026-09-29)
41
42 * **6 steps, a different balance.** Against v0.2.1: less grain and cleaner surfaces, fine texture a little softer;
43 diversity (still close to the base model's) and small-text accuracy about the same. It is not a strict upgrade: if
44 you prefer the crisper look, v0.2.1 is still in the repository.
45 * **New 9-step mode:** 7 turbo steps, then the LoRA is switched off and the base model finishes the last two.
46 Finer detail, and small text comes out right more often (not always). It takes about 1.4–1.5× as long as 6 steps
47 (still about 3.5× faster than the base model). It works in diffusers and the demo Space only ([9 steps](#9-steps)).
48 * **We think 6 steps is close to its capacity.** Since v0.2.1, every gain we found at 6 steps cost something
49 elsewhere: sharper came with more grain, less grain came with a softer look. Beyond this point, quality most likely
50 has to be paid for with steps, which is what the 9-step mode does.
51
52 | file | |
53 |---|---|
54 | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors` | **v0.3 LoRA** (rank 256, bf16, 1.3 GB), loaded on the base transformer at runtime. **Use this.** |
55 | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r128.safetensors` | the same adapter cut to rank 128 (0.7 GB), used by the [ComfyUI workflows](#comfyui) |
56 | `peft_v0.3/` | the v0.3 adapter in peft key format |
57 | `...-{v0.3,v0.2.1}-6step-{int8_convrot,fp8_e4m3fn}.safetensors`, `...-{Q8_0,Q6_K,Q5_K_M,Q4_K_M}.gguf` | the LoRA merged into the base transformer, one file per format, for ComfyUI ([single-file transformers](#single-file-transformers)) |
58 | `comfyui/` | ComfyUI custom nodes, text-to-image / edit workflows, example inputs |
59 | `scheduler/` | the base scheduler config with `shift_terminal: null` |
60 | `...-v0.2.1-6step-lora-r256/r128.safetensors`, `peft_v0.2.1/` | v0.2.1 (2026-09-24), same usage |
61 | `...-v0.2-5step-lora-r256/r128.safetensors`, `peft_v0.2/` | v0.2 (2026-09-23), same usage (6 steps) |
62
63 ## Install
64
65 ```bash
66 pip install -U torch "transformers>=5.17,<6" accelerate safetensors peft pillow
67 pip install "git+https://github.com/huggingface/diffusers.git@80c7ed262aeffbeb43ef13ae04baeb9b84515a69"
68 ```
69
70 `QwenImage21Pipeline` is not in a released `diffusers` yet, hence the pinned git install. `peft` is required.
71
72 ## Usage
73
74 ```python
75 import torch
76 from diffusers import QwenImage21Pipeline, FlowMatchEulerDiscreteScheduler
77
78 pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16)
79 pipe.load_lora_weights(
80 "Viggle/Qwen-Image-2.1-viggle-turbo",
81 weight_name="Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors",
82 )
83 pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
84 "Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="scheduler"
85 )
86 pipe.to("cuda")
87
88 STEPS, SIGMAS = 6, [1.0, 0.9375, 0.875, 0.75, 0.5, 0.25] # pass both to every call
89 ```
90
91 ### Text to image
92
93 ```python
94 image = pipe(
95 prompt="A studio portrait of an old fisherman mending a net, warm rim light, 85mm.",
96 height=1024,
97 width=1024,
98 num_inference_steps=STEPS,
99 sigmas=SIGMAS,
100 true_cfg_scale=1.0, # no CFG (also the default)
101 generator=torch.Generator("cuda").manual_seed(0),
102 ).images[0]
103 image.save("out.png")
104 ```
105
106 ### Image editing (1–3 reference images)
107
108 ```python
109 from diffusers.utils import load_image
110
111 image = pipe( # same pipe object as above
112 prompt="Replace the background with a sunset beach, keep the subject unchanged.",
113 image=[load_image("input.png")], # list; order fixes <image1>, <image2>, ...
114 output_resolution=1024,
115 num_inference_steps=STEPS,
116 sigmas=SIGMAS,
117 true_cfg_scale=1.0,
118 generator=torch.Generator("cuda").manual_seed(0),
119 ).images[0]
120 ```
121
122 ### 9 steps
123
124 ```python
125 SIGMAS_9 = [1.0, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25, 1 / 6, 1 / 12]
126
127 # The pipeline computes the text/reference K/V once and reuses them; those come from the turbo, so the first
128 # base-model step computes them again.
129 forward = pipe.transformer.forward
130 reextract = [False]
131
132 def transformer_forward(*args, kv_cache_mode=None, **kwargs):
133 if reextract[0] and kv_cache_mode == "cached":
134 kv_cache_mode, reextract[0] = "extract", False
135 return forward(*args, kv_cache_mode=kv_cache_mode, **kwargs)
136
137 pipe.transformer.forward = transformer_forward
138
139 def base_tail(pipe, i, t, kwargs):
140 if i == 6: # runs after the 7th step; the base model takes the last two
141 pipe.disable_lora()
142 reextract[0] = True
143 return kwargs
144
145 image = pipe(
146 prompt="...", # works for editing too
147 num_inference_steps=9,
148 sigmas=SIGMAS_9,
149 true_cfg_scale=1.0,
150 callback_on_step_end=base_tail,
151 generator=torch.Generator("cuda").manual_seed(0),
152 ).images[0]
153 pipe.enable_lora() # back to the turbo for the next call
154 ```
155
156 ### Rules that matter
157
158 * **6 steps with `sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25]`, `true_cfg_scale=1.0`, no negative prompt.** These
159 are raw nodes: the pipeline applies its resolution-dependent shift to them, so pass them as written at every size.
160 * **To change the step count, add or remove steps at the high-noise end only**, and keep `0.875, 0.75, 0.5, 0.25`:
161 5 steps `[1, 0.875, 0.75, 0.5, 0.25]`, 7 steps `[1, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25]`. Moving the low-noise
162 nodes makes images softer; plain `num_inference_steps` without `sigmas=` and CFG do not help.
163 * **Use the shipped scheduler config.** The base config's `shift_terminal: 0.02` wrecks the last step.
164 * **Keep the LoRA unmerged and at scale 1.0.** Merging it into bf16 weights loses part of the update. For ComfyUI there
165 are also [single-file transformers](#single-file-transformers) with the LoRA merged in fp32, then quantized once.
166 * Reference order decides which image `image 1` / `image 2` in the prompt refers to. Without `height`/`width`, the
167 output aspect ratio follows the **last** reference (in ComfyUI: the **first**).
168 * Prompt rewriting with the official [PE-T2I](https://huggingface.co/Qwen/Qwen-Image-2.1-PE-T2I) /
169 [PE-I2I](https://huggingface.co/Qwen/Qwen-Image-2.1-PE-I2I) rewriters helps composition and rendered text. Raw
170 prompts work too.
171 * About 1 MP is the sweet spot; up to about 4 MP works. Keep width and height at multiples of 16.
172 * `peft` users can load `peft_v0.3/` instead:
173 `pipe.transformer.load_lora_adapter("Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="peft_v0.3", weight_name="adapter_model.safetensors", prefix=None)`.
174 It holds the same weights under different key names, so pick one of the two, not both.
175
176 ### Rank 128
177
178 The r128 file is the r256 adapter truncated by an exact per-layer SVD of its update. What the cut drops is about ten
179 times smaller than bf16 rounding of the base weights.
180
181 ## ComfyUI
182
183 > The ComfyUI port is mostly vibe-coded with an AI coding assistant, and I don't use ComfyUI day to day. The sigma
184 > schedule matches diffusers to float precision and the workflows run end to end, but expect rough edges. Issues
185 > and fixes are very welcome.
186
187 Tested with ComfyUI 0.37.0 (frontend 1.53.6), which has native Qwen-Image-2.1 support. In [`comfyui/`](comfyui):
188
189 * `viggle_turbo.py` — two custom nodes. Copy it into `ComfyUI/custom_nodes/` and restart ComfyUI.
190 * `Qwen-Image-2.1-viggle-turbo-t2i.json`, `Qwen-Image-2.1-viggle-turbo-edit.json` — the workflows (drag into ComfyUI).
191 * `input/woman2.webp`, `input/cat.webp` — the edit workflow's example references (from the
192 [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space); copy
193 them into `ComfyUI/input/`.
194
195 | ComfyUI folder | file | size |
196 |---|---|---|
197 | `diffusion_models/` | [`qwen_image_2.1_int8_convrot.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) or `qwen_image_2.1_bf16.safetensors` | 7.3 / 14.2 GB |
198 | `text_encoders/` | [`qwen3vl_8b_int8_convrot.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) or `qwen3vl_8b_bf16.safetensors` | 9.4 / 17.5 GB |
199 | `vae/` | [`qwen_image_2.1_vae_bf16.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) | 0.7 GB |
200 | `loras/` | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r128.safetensors` (this repo) or the r256 file | 0.7 / 1.3 GB |
201
202 The workflows default to the int8 files, the r128 LoRA and the prompt enhancer on; this peaks at about 26 GB of VRAM
203 at 1248 × 832.
204
205 * **Viggle Turbo Sigmas** — the 6-step schedule with the pipeline's resolution-dependent shift. Use it instead of a
206 KSampler scheduler, with euler and `BasicGuider` (no CFG, no negative prompt).
207 * **Viggle Turbo LoRA (unmerged)** — applies the LoRA at runtime, as diffusers does. The stock LoRA loaders merge it
208 into the weights, which drops about 30% of this adapter's update on bf16 and adds noise on int8. The unmerged node
209 costs 10–25% more time per step. Keep its strength at 1.0.
210
211 The 9-step mode is not in the workflows. In the edit workflow the output size follows **image 1**, at about 1 MP.
212
213 **Troubleshooting:** with `comfy_kitchen` 0.2.35 on an NVIDIA driver older than 580, `TextGenerate` (the prompt
214 enhancer) fails with a CUDA driver error. Update the driver, or turn *Enhance prompt* off.
215
216 ### Single-file transformers
217
218 If you would rather not load a LoRA, these files are the base transformer with the r256 LoRA already merged in
219 (in fp32, then quantized once). Put the file in `models/diffusion_models/` (`.gguf` files need the
220 [ComfyUI-GGUF](https://github.com/city96/ComfyUI-GGUF) nodes) and open the matching workflow in `comfyui/`. The text
221 encoder and VAE are the same Comfy-Org files as above, and you still need **Viggle Turbo Sigmas** from
222 `comfyui/viggle_turbo.py`. Only the 6-step mode is available as a single file.
223
224 | format | file suffix | size | ComfyUI loader | LPIPS ¹ v0.3 | LPIPS ¹ v0.2.1 |
225 |---|---|---|---|---|---|
226 | GGUF Q8_0 | `-6step-Q8_0.gguf` | 7.7 GB | Unet Loader (GGUF) | 0.051 | 0.054 |
227 | int8 (Comfy-Org convrot recipe) | `-6step-int8_convrot.safetensors` | 7.3 GB | Load Diffusion Model | 0.057 | 0.060 |
228 | GGUF Q6_K | `-6step-Q6_K.gguf` | 6.0 GB | Unet Loader (GGUF) | 0.055 | 0.067 |
229 | fp8 e4m3fn, weight-only | `-6step-fp8_e4m3fn.safetensors` | 7.3 GB | Load Diffusion Model | 0.068 | 0.070 |
230 | GGUF Q5_K_M | `-6step-Q5_K_M.gguf` | 5.1 GB | Unet Loader (GGUF) | 0.076 | 0.083 |
231 | GGUF Q4_K_M ² | `-6step-Q4_K_M.gguf` | 4.3 GB | Unet Loader (GGUF) | 0.100 | 0.118 |
232 | *reference:* Comfy-Org int8 + LoRA r128 (the LoRA workflows) | | 7.3 + 0.7 GB | | 0.041 | 0.044 |
233
234 Full names are `Qwen-Image-2.1-viggle-turbo-v0.3` or `-v0.2.1` followed by the suffix. Workflows:
235 `comfyui/Qwen-Image-2.1-viggle-turbo-{v0.3,v0.2.1}-6step-merged-{t2i,edit}.json` (int8 by default; pick the fp8 file
236 in the loader) and `…-6step-gguf-{t2i,edit}.json` (Q8_0 by default; pick another GGUF in the loader).
237
238 **Merged is close to, but not the same as, the LoRA path.** On most requests the image matches the LoRA workflow up to
239 fine detail, but on a few (about 8 in 96 at int8 or Q8_0, against 3–5 in 96 for the LoRA workflow) the merged model
240 settles on a different composition or outfit. The result is not necessarily worse, just different. If you need
241 outputs that match diffusers, use the LoRA workflows.
242
243 ² **Q4_K_M drifts visibly more** (23–29 of 96 requests change noticeably). Use it only if nothing larger fits in memory.
244
245 ¹ mean LPIPS (VGG, ≤512 px) against diffusers with the r256 LoRA on our 96 held-out requests, same prompt, inputs,
246 seed and noise; lower is closer. For scale, ComfyUI and diffusers differ by about 0.03–0.04 with no LoRA at all.
247
248 ## Known limitations
249
250 * **Complicated edits** (multi-reference composition, face swaps, identity-preserving edits, instructions with
251 several constraints) can still fall short of the base model: duplicated or ghosted figures, identity drift.
252 * **Small or long rendered text** garbles more often than with the base model. 9 steps often helps, not always.
253 * Colours come out a few percent less saturated than the base model's.
254 * At 6 steps, v0.3 is a little softer on fine texture than v0.2.1.
255 * 2K output, RGBA output, mask-guided edits and edits with more than 3 references are checked only by eye on the
256 Comparison tab examples. No standard benchmark is claimed.
257
258 ## License
259
260 This model is a derivative work of Qwen-Image-2.1 and is distributed under the **Qwen RESEARCH LICENSE AGREEMENT**
261 ([`LICENSE`](LICENSE)): **non-commercial use only** — research or evaluation purposes. Commercial use requires a
262 separate licence from the licensor (`model-business@notice.qwencloud.com`). See [`NOTICE`](NOTICE) for the required
263 attribution.
264
265 > Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology
266 > Co., Ltd. All Rights Reserved.
267
268 Relative to [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) this repository **adds** LoRA
269 adapters (v0.3, v0.2.1 and v0.2, at rank 256 and 128), single-file transformers (the base transformer with the v0.3 or
270 v0.2.1 adapter merged in, quantized to int8, fp8 or GGUF), ComfyUI nodes, workflows and two example input photos, and a
271 scheduler config with `shift_terminal` changed from `0.02` to `null`. Text encoder, VAE and processor are not
272 redistributed; the base transformer only in modified form (see [`NOTICE`](NOTICE)).
273
274 Distillation and release by **Viggle**. **Built with Qwen.**
275