README.md
| 1 | --- |
| 2 | license: other |
| 3 | license_name: qwen-research |
| 4 | license_link: LICENSE |
| 5 | base_model: Qwen/Qwen-Image-2.1 |
| 6 | base_model_relation: adapter |
| 7 | library_name: diffusers |
| 8 | pipeline_tag: text-to-image |
| 9 | tags: |
| 10 | - diffusers |
| 11 | - lora |
| 12 | - text-to-image |
| 13 | - image-to-image |
| 14 | - image-editing |
| 15 | - distillation |
| 16 | - turbo |
| 17 | - few-step |
| 18 | - qwen-image |
| 19 | - comfyui |
| 20 | - gguf |
| 21 | - int8 |
| 22 | - fp8 |
| 23 | - quantized |
| 24 | --- |
| 25 | |
| 26 | # Qwen-Image-2.1-viggle-turbo — v0.3 |
| 27 | |
| 28 | **Built with Qwen.** A few-step distilled version of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1) |
| 29 | by Viggle. Text-to-image and instruction-driven editing with 1–3 reference images in **6 steps instead of 40**, with |
| 30 | **no classifier-free guidance**. |
| 31 | |
| 32 | <video controls autoplay muted loop playsinline width="100%" |
| 33 | src="https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo/resolve/main/assets/viggle_turbo_promo.mp4"></video> |
| 34 | |
| 35 | About **5× faster** than the 40-step base model end to end, and very competitive with it in quality: on the official |
| 36 | Qwen examples the two are hard to tell apart on most prompts. The clearest gaps are small, dense text and complicated |
| 37 | edits ([Known limitations](#known-limitations)). Compare them yourself in the **Comparison** tab of the |
| 38 | [demo Space](https://huggingface.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo). |
| 39 | |
| 40 | ## v0.3 (2026-09-29) |
| 41 | |
| 42 | * **6 steps, a different balance.** Against v0.2.1: less grain and cleaner surfaces, fine texture a little softer; |
| 43 | diversity (still close to the base model's) and small-text accuracy about the same. It is not a strict upgrade: if |
| 44 | you prefer the crisper look, v0.2.1 is still in the repository. |
| 45 | * **New 9-step mode:** 7 turbo steps, then the LoRA is switched off and the base model finishes the last two. |
| 46 | Finer detail, and small text comes out right more often (not always). It takes about 1.4–1.5× as long as 6 steps |
| 47 | (still about 3.5× faster than the base model). It works in diffusers and the demo Space only ([9 steps](#9-steps)). |
| 48 | * **We think 6 steps is close to its capacity.** Since v0.2.1, every gain we found at 6 steps cost something |
| 49 | elsewhere: sharper came with more grain, less grain came with a softer look. Beyond this point, quality most likely |
| 50 | has to be paid for with steps, which is what the 9-step mode does. |
| 51 | |
| 52 | | file | | |
| 53 | |---|---| |
| 54 | | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors` | **v0.3 LoRA** (rank 256, bf16, 1.3 GB), loaded on the base transformer at runtime. **Use this.** | |
| 55 | | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r128.safetensors` | the same adapter cut to rank 128 (0.7 GB), used by the [ComfyUI workflows](#comfyui) | |
| 56 | | `peft_v0.3/` | the v0.3 adapter in peft key format | |
| 57 | | `...-{v0.3,v0.2.1}-6step-{int8_convrot,fp8_e4m3fn}.safetensors`, `...-{Q8_0,Q6_K,Q5_K_M,Q4_K_M}.gguf` | the LoRA merged into the base transformer, one file per format, for ComfyUI ([single-file transformers](#single-file-transformers)) | |
| 58 | | `comfyui/` | ComfyUI custom nodes, text-to-image / edit workflows, example inputs | |
| 59 | | `scheduler/` | the base scheduler config with `shift_terminal: null` | |
| 60 | | `...-v0.2.1-6step-lora-r256/r128.safetensors`, `peft_v0.2.1/` | v0.2.1 (2026-09-24), same usage | |
| 61 | | `...-v0.2-5step-lora-r256/r128.safetensors`, `peft_v0.2/` | v0.2 (2026-09-23), same usage (6 steps) | |
| 62 | |
| 63 | ## Install |
| 64 | |
| 65 | ```bash |
| 66 | pip install -U torch "transformers>=5.17,<6" accelerate safetensors peft pillow |
| 67 | pip install "git+https://github.com/huggingface/diffusers.git@80c7ed262aeffbeb43ef13ae04baeb9b84515a69" |
| 68 | ``` |
| 69 | |
| 70 | `QwenImage21Pipeline` is not in a released `diffusers` yet, hence the pinned git install. `peft` is required. |
| 71 | |
| 72 | ## Usage |
| 73 | |
| 74 | ```python |
| 75 | import torch |
| 76 | from diffusers import QwenImage21Pipeline, FlowMatchEulerDiscreteScheduler |
| 77 | |
| 78 | pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16) |
| 79 | pipe.load_lora_weights( |
| 80 | "Viggle/Qwen-Image-2.1-viggle-turbo", |
| 81 | weight_name="Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r256.safetensors", |
| 82 | ) |
| 83 | pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained( |
| 84 | "Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="scheduler" |
| 85 | ) |
| 86 | pipe.to("cuda") |
| 87 | |
| 88 | STEPS, SIGMAS = 6, [1.0, 0.9375, 0.875, 0.75, 0.5, 0.25] # pass both to every call |
| 89 | ``` |
| 90 | |
| 91 | ### Text to image |
| 92 | |
| 93 | ```python |
| 94 | image = pipe( |
| 95 | prompt="A studio portrait of an old fisherman mending a net, warm rim light, 85mm.", |
| 96 | height=1024, |
| 97 | width=1024, |
| 98 | num_inference_steps=STEPS, |
| 99 | sigmas=SIGMAS, |
| 100 | true_cfg_scale=1.0, # no CFG (also the default) |
| 101 | generator=torch.Generator("cuda").manual_seed(0), |
| 102 | ).images[0] |
| 103 | image.save("out.png") |
| 104 | ``` |
| 105 | |
| 106 | ### Image editing (1–3 reference images) |
| 107 | |
| 108 | ```python |
| 109 | from diffusers.utils import load_image |
| 110 | |
| 111 | image = pipe( # same pipe object as above |
| 112 | prompt="Replace the background with a sunset beach, keep the subject unchanged.", |
| 113 | image=[load_image("input.png")], # list; order fixes <image1>, <image2>, ... |
| 114 | output_resolution=1024, |
| 115 | num_inference_steps=STEPS, |
| 116 | sigmas=SIGMAS, |
| 117 | true_cfg_scale=1.0, |
| 118 | generator=torch.Generator("cuda").manual_seed(0), |
| 119 | ).images[0] |
| 120 | ``` |
| 121 | |
| 122 | ### 9 steps |
| 123 | |
| 124 | ```python |
| 125 | SIGMAS_9 = [1.0, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25, 1 / 6, 1 / 12] |
| 126 | |
| 127 | # The pipeline computes the text/reference K/V once and reuses them; those come from the turbo, so the first |
| 128 | # base-model step computes them again. |
| 129 | forward = pipe.transformer.forward |
| 130 | reextract = [False] |
| 131 | |
| 132 | def transformer_forward(*args, kv_cache_mode=None, **kwargs): |
| 133 | if reextract[0] and kv_cache_mode == "cached": |
| 134 | kv_cache_mode, reextract[0] = "extract", False |
| 135 | return forward(*args, kv_cache_mode=kv_cache_mode, **kwargs) |
| 136 | |
| 137 | pipe.transformer.forward = transformer_forward |
| 138 | |
| 139 | def base_tail(pipe, i, t, kwargs): |
| 140 | if i == 6: # runs after the 7th step; the base model takes the last two |
| 141 | pipe.disable_lora() |
| 142 | reextract[0] = True |
| 143 | return kwargs |
| 144 | |
| 145 | image = pipe( |
| 146 | prompt="...", # works for editing too |
| 147 | num_inference_steps=9, |
| 148 | sigmas=SIGMAS_9, |
| 149 | true_cfg_scale=1.0, |
| 150 | callback_on_step_end=base_tail, |
| 151 | generator=torch.Generator("cuda").manual_seed(0), |
| 152 | ).images[0] |
| 153 | pipe.enable_lora() # back to the turbo for the next call |
| 154 | ``` |
| 155 | |
| 156 | ### Rules that matter |
| 157 | |
| 158 | * **6 steps with `sigmas=[1.0, 0.9375, 0.875, 0.75, 0.5, 0.25]`, `true_cfg_scale=1.0`, no negative prompt.** These |
| 159 | are raw nodes: the pipeline applies its resolution-dependent shift to them, so pass them as written at every size. |
| 160 | * **To change the step count, add or remove steps at the high-noise end only**, and keep `0.875, 0.75, 0.5, 0.25`: |
| 161 | 5 steps `[1, 0.875, 0.75, 0.5, 0.25]`, 7 steps `[1, 0.9583, 0.9167, 0.875, 0.75, 0.5, 0.25]`. Moving the low-noise |
| 162 | nodes makes images softer; plain `num_inference_steps` without `sigmas=` and CFG do not help. |
| 163 | * **Use the shipped scheduler config.** The base config's `shift_terminal: 0.02` wrecks the last step. |
| 164 | * **Keep the LoRA unmerged and at scale 1.0.** Merging it into bf16 weights loses part of the update. For ComfyUI there |
| 165 | are also [single-file transformers](#single-file-transformers) with the LoRA merged in fp32, then quantized once. |
| 166 | * Reference order decides which image `image 1` / `image 2` in the prompt refers to. Without `height`/`width`, the |
| 167 | output aspect ratio follows the **last** reference (in ComfyUI: the **first**). |
| 168 | * Prompt rewriting with the official [PE-T2I](https://huggingface.co/Qwen/Qwen-Image-2.1-PE-T2I) / |
| 169 | [PE-I2I](https://huggingface.co/Qwen/Qwen-Image-2.1-PE-I2I) rewriters helps composition and rendered text. Raw |
| 170 | prompts work too. |
| 171 | * About 1 MP is the sweet spot; up to about 4 MP works. Keep width and height at multiples of 16. |
| 172 | * `peft` users can load `peft_v0.3/` instead: |
| 173 | `pipe.transformer.load_lora_adapter("Viggle/Qwen-Image-2.1-viggle-turbo", subfolder="peft_v0.3", weight_name="adapter_model.safetensors", prefix=None)`. |
| 174 | It holds the same weights under different key names, so pick one of the two, not both. |
| 175 | |
| 176 | ### Rank 128 |
| 177 | |
| 178 | The r128 file is the r256 adapter truncated by an exact per-layer SVD of its update. What the cut drops is about ten |
| 179 | times smaller than bf16 rounding of the base weights. |
| 180 | |
| 181 | ## ComfyUI |
| 182 | |
| 183 | > The ComfyUI port is mostly vibe-coded with an AI coding assistant, and I don't use ComfyUI day to day. The sigma |
| 184 | > schedule matches diffusers to float precision and the workflows run end to end, but expect rough edges. Issues |
| 185 | > and fixes are very welcome. |
| 186 | |
| 187 | Tested with ComfyUI 0.37.0 (frontend 1.53.6), which has native Qwen-Image-2.1 support. In [`comfyui/`](comfyui): |
| 188 | |
| 189 | * `viggle_turbo.py` — two custom nodes. Copy it into `ComfyUI/custom_nodes/` and restart ComfyUI. |
| 190 | * `Qwen-Image-2.1-viggle-turbo-t2i.json`, `Qwen-Image-2.1-viggle-turbo-edit.json` — the workflows (drag into ComfyUI). |
| 191 | * `input/woman2.webp`, `input/cat.webp` — the edit workflow's example references (from the |
| 192 | [black-forest-labs/flux-klein-9b-kv](https://huggingface.co/spaces/black-forest-labs/flux-klein-9b-kv) Space); copy |
| 193 | them into `ComfyUI/input/`. |
| 194 | |
| 195 | | ComfyUI folder | file | size | |
| 196 | |---|---|---| |
| 197 | | `diffusion_models/` | [`qwen_image_2.1_int8_convrot.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) or `qwen_image_2.1_bf16.safetensors` | 7.3 / 14.2 GB | |
| 198 | | `text_encoders/` | [`qwen3vl_8b_int8_convrot.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) or `qwen3vl_8b_bf16.safetensors` | 9.4 / 17.5 GB | |
| 199 | | `vae/` | [`qwen_image_2.1_vae_bf16.safetensors`](https://huggingface.co/Comfy-Org/Qwen-Image-2.1) | 0.7 GB | |
| 200 | | `loras/` | `Qwen-Image-2.1-viggle-turbo-v0.3-6step-lora-r128.safetensors` (this repo) or the r256 file | 0.7 / 1.3 GB | |
| 201 | |
| 202 | The workflows default to the int8 files, the r128 LoRA and the prompt enhancer on; this peaks at about 26 GB of VRAM |
| 203 | at 1248 × 832. |
| 204 | |
| 205 | * **Viggle Turbo Sigmas** — the 6-step schedule with the pipeline's resolution-dependent shift. Use it instead of a |
| 206 | KSampler scheduler, with euler and `BasicGuider` (no CFG, no negative prompt). |
| 207 | * **Viggle Turbo LoRA (unmerged)** — applies the LoRA at runtime, as diffusers does. The stock LoRA loaders merge it |
| 208 | into the weights, which drops about 30% of this adapter's update on bf16 and adds noise on int8. The unmerged node |
| 209 | costs 10–25% more time per step. Keep its strength at 1.0. |
| 210 | |
| 211 | The 9-step mode is not in the workflows. In the edit workflow the output size follows **image 1**, at about 1 MP. |
| 212 | |
| 213 | **Troubleshooting:** with `comfy_kitchen` 0.2.35 on an NVIDIA driver older than 580, `TextGenerate` (the prompt |
| 214 | enhancer) fails with a CUDA driver error. Update the driver, or turn *Enhance prompt* off. |
| 215 | |
| 216 | ### Single-file transformers |
| 217 | |
| 218 | If you would rather not load a LoRA, these files are the base transformer with the r256 LoRA already merged in |
| 219 | (in fp32, then quantized once). Put the file in `models/diffusion_models/` (`.gguf` files need the |
| 220 | [ComfyUI-GGUF](https://github.com/city96/ComfyUI-GGUF) nodes) and open the matching workflow in `comfyui/`. The text |
| 221 | encoder and VAE are the same Comfy-Org files as above, and you still need **Viggle Turbo Sigmas** from |
| 222 | `comfyui/viggle_turbo.py`. Only the 6-step mode is available as a single file. |
| 223 | |
| 224 | | format | file suffix | size | ComfyUI loader | LPIPS ¹ v0.3 | LPIPS ¹ v0.2.1 | |
| 225 | |---|---|---|---|---|---| |
| 226 | | GGUF Q8_0 | `-6step-Q8_0.gguf` | 7.7 GB | Unet Loader (GGUF) | 0.051 | 0.054 | |
| 227 | | int8 (Comfy-Org convrot recipe) | `-6step-int8_convrot.safetensors` | 7.3 GB | Load Diffusion Model | 0.057 | 0.060 | |
| 228 | | GGUF Q6_K | `-6step-Q6_K.gguf` | 6.0 GB | Unet Loader (GGUF) | 0.055 | 0.067 | |
| 229 | | fp8 e4m3fn, weight-only | `-6step-fp8_e4m3fn.safetensors` | 7.3 GB | Load Diffusion Model | 0.068 | 0.070 | |
| 230 | | GGUF Q5_K_M | `-6step-Q5_K_M.gguf` | 5.1 GB | Unet Loader (GGUF) | 0.076 | 0.083 | |
| 231 | | GGUF Q4_K_M ² | `-6step-Q4_K_M.gguf` | 4.3 GB | Unet Loader (GGUF) | 0.100 | 0.118 | |
| 232 | | *reference:* Comfy-Org int8 + LoRA r128 (the LoRA workflows) | | 7.3 + 0.7 GB | | 0.041 | 0.044 | |
| 233 | |
| 234 | Full names are `Qwen-Image-2.1-viggle-turbo-v0.3` or `-v0.2.1` followed by the suffix. Workflows: |
| 235 | `comfyui/Qwen-Image-2.1-viggle-turbo-{v0.3,v0.2.1}-6step-merged-{t2i,edit}.json` (int8 by default; pick the fp8 file |
| 236 | in the loader) and `…-6step-gguf-{t2i,edit}.json` (Q8_0 by default; pick another GGUF in the loader). |
| 237 | |
| 238 | **Merged is close to, but not the same as, the LoRA path.** On most requests the image matches the LoRA workflow up to |
| 239 | fine detail, but on a few (about 8 in 96 at int8 or Q8_0, against 3–5 in 96 for the LoRA workflow) the merged model |
| 240 | settles on a different composition or outfit. The result is not necessarily worse, just different. If you need |
| 241 | outputs that match diffusers, use the LoRA workflows. |
| 242 | |
| 243 | ² **Q4_K_M drifts visibly more** (23–29 of 96 requests change noticeably). Use it only if nothing larger fits in memory. |
| 244 | |
| 245 | ¹ mean LPIPS (VGG, ≤512 px) against diffusers with the r256 LoRA on our 96 held-out requests, same prompt, inputs, |
| 246 | seed and noise; lower is closer. For scale, ComfyUI and diffusers differ by about 0.03–0.04 with no LoRA at all. |
| 247 | |
| 248 | ## Known limitations |
| 249 | |
| 250 | * **Complicated edits** (multi-reference composition, face swaps, identity-preserving edits, instructions with |
| 251 | several constraints) can still fall short of the base model: duplicated or ghosted figures, identity drift. |
| 252 | * **Small or long rendered text** garbles more often than with the base model. 9 steps often helps, not always. |
| 253 | * Colours come out a few percent less saturated than the base model's. |
| 254 | * At 6 steps, v0.3 is a little softer on fine texture than v0.2.1. |
| 255 | * 2K output, RGBA output, mask-guided edits and edits with more than 3 references are checked only by eye on the |
| 256 | Comparison tab examples. No standard benchmark is claimed. |
| 257 | |
| 258 | ## License |
| 259 | |
| 260 | This model is a derivative work of Qwen-Image-2.1 and is distributed under the **Qwen RESEARCH LICENSE AGREEMENT** |
| 261 | ([`LICENSE`](LICENSE)): **non-commercial use only** — research or evaluation purposes. Commercial use requires a |
| 262 | separate licence from the licensor (`model-business@notice.qwencloud.com`). See [`NOTICE`](NOTICE) for the required |
| 263 | attribution. |
| 264 | |
| 265 | > Qwen is licensed under the Qwen RESEARCH LICENSE AGREEMENT, Copyright (c) 2026 Hangzhou Tongyi Laboratory Technology |
| 266 | > Co., Ltd. All Rights Reserved. |
| 267 | |
| 268 | Relative to [`Qwen/Qwen-Image-2.1`](https://huggingface.co/Qwen/Qwen-Image-2.1) this repository **adds** LoRA |
| 269 | adapters (v0.3, v0.2.1 and v0.2, at rank 256 and 128), single-file transformers (the base transformer with the v0.3 or |
| 270 | v0.2.1 adapter merged in, quantized to int8, fp8 or GGUF), ComfyUI nodes, workflows and two example input photos, and a |
| 271 | scheduler config with `shift_terminal` changed from `0.02` to `null`. Text encoder, VAE and processor are not |
| 272 | redistributed; the base transformer only in modified form (see [`NOTICE`](NOTICE)). |
| 273 | |
| 274 | Distillation and release by **Viggle**. **Built with Qwen.** |
| 275 | |