README.md
| 1 | --- |
| 2 | base_model: Qwen/Qwen-Image-2.1 |
| 3 | base_model_relation: quantized |
| 4 | license: other |
| 5 | license_name: qwen-research |
| 6 | license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE |
| 7 | language: |
| 8 | - en |
| 9 | - zh |
| 10 | pipeline_tag: text-to-image |
| 11 | tags: |
| 12 | - gguf |
| 13 | - quantized |
| 14 | - unsloth |
| 15 | - qwen |
| 16 | - image-generation |
| 17 | widget: |
| 18 | - text: Photorealistic editorial photograph of a woman barista making a latte in a modern, minimalist café on a sunny tropical morning. |
| 19 | output: |
| 20 | url: assets/cafe.png |
| 21 | - text: A lone astronaut crossing a dark, frozen lake beneath enormous rings stretching across an alien sky. Fine cracks visible under the translucent ice, distant mountains, soft blue twilight, cinematic scale, photorealistic detail. |
| 22 | output: |
| 23 | url: assets/spaces.png |
| 24 | --- |
| 25 | # Read our How to [Run Qwen-Image-2.1 Guide!](https://unsloth.ai/docs/models/qwen-image-2.1) 💜 |
| 26 | |
| 27 | This is a GGUF quantized version of [Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1). <br> |
| 28 | unsloth/Qwen-Image-2.1-GGUF uses [Unsloth Dynamic 2.0](https://docs.unsloth.ai/basics/unsloth-dynamic-2.0-ggufs) methodology for SOTA performance. |
| 29 | |
| 30 | - Important layers are upcasted to higher precision, per tensor, from a measured sensitivity scan. |
| 31 | - Run these with [Unsloth Desktop](https://github.com/unslothai/unsloth), stable-diffusion.cpp and more. A GGUF is the denoiser only, so it needs the VAE and the Qwen3-VL text encoder alongside it. |
| 32 | - VAE: [unsloth/Qwen-Image-2.1-FP8](https://huggingface.co/unsloth/Qwen-Image-2.1-FP8) `vae/qwen_image_2.1_vae_bf16.safetensors`. Text encoder: [unsloth/Qwen3-VL-8B-Instruct-GGUF](https://huggingface.co/unsloth/Qwen3-VL-8B-Instruct-GGUF) `Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf`, the Dynamic 2.0 4-bit rung rather than the uniform `Q4_K_M`. Measured against the `Q4_K_M` encoder at a shared seed, with the denoiser and VAE held fixed: LPIPS 0.029, SSIM 0.959, 5.15 GB vs 5.03 GB, 36.5 s vs 39.0 s. |
| 33 | <div> |
| 34 | <div style="display: flex; gap: 5px; align-items: center; "> |
| 35 | <a href="https://github.com/unslothai/unsloth/"> |
| 36 | <img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="133"> |
| 37 | </a> |
| 38 | <a href="https://discord.gg/unsloth"> |
| 39 | <img src="https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png" width="173"> |
| 40 | </a> |
| 41 | <a href="https://unsloth.ai/docs/models/qwen-image-2.1"> |
| 42 | <img src="https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png" width="143"> |
| 43 | </a> |
| 44 | </div> |
| 45 | </div> |
| 46 | See below for image editing operating inside of Unsloth Desktop: |
| 47 | <img width="600" alt="qwen-image-2.1 unsloth desktop" src="https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8E4zaQZmeYBhYIXkvveM%252F01-2edit-multi-image.webp%3Falt%3Dmedia%26token%3Dad7d102c-b304-4c75-b66d-4a68cd518a37&width=768&dpr=3&quality=100&sign=a5a37f593c66ddc27c847bab0d138eda&sv=3" /> |
| 48 | |
| 49 | |
| 50 | ```bash |
| 51 | sd-cli --diffusion-model qwen-image-2.1-Q4_K_M.gguf \ |
| 52 | --vae qwen_image_2.1_vae_bf16.safetensors \ |
| 53 | --llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf \ |
| 54 | -p "a cartoon sloth mascot waving, flat vector illustration, bright colours" \ |
| 55 | --steps 20 --cfg-scale 6.0 --sampling-method euler -W 1024 -H 1024 --diffusion-fa \ |
| 56 | -o out.png |
| 57 | ``` |
| 58 | |
| 59 | ### Samples |
| 60 | |
| 61 | Rendered with the Q4_K_M denoiser and the Q4_K_M text encoder, 1024x1024, 20 steps, cfg 6.0, euler. |
| 62 | |
| 63 | <table> |
| 64 | <tr> |
| 65 | <td><img src="assets/cafe.png" width="200"></td> |
| 66 | <td><img src="assets/spaces.png" width="200"></td> |
| 67 | </tr> |
| 68 | <tr> |
| 69 | <td><img src="assets/cinema.png" width="200"></td> |
| 70 | <td><img src="assets/cardesert.png" width="200"></td> |
| 71 | </tr> |
| 72 | </table> |
| 73 | |
| 74 | --- |
| 75 | |
| 76 | <p align="center"> |
| 77 | <img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-Image/image2.1/logo.png" width="400"/> |
| 78 | </p> |
| 79 | <p align="center"> |
| 80 | 🤖 <a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">ModelScope</a> | |
| 81 | 🤗 <a href="https://huggingface.co/Qwen/Qwen-Image-2.1">HuggingFace</a> | |
| 82 | 📑 <a href="https://qwen.ai/blog?id=qwen-image-2.1">Blog</a> | |
| 83 | 🖥️ <a href="https://huggingface.co/spaces/Qwen/Qwen-Image-2.1">Demo</a> | |
| 84 | 🫨 <a href="https://discord.gg/BEYSk3pkSu">Discord</a> | |
| 85 | 💬 <a href="https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/assets/qr.png">WeChat</a> |
| 86 | </p> |
| 87 | |
| 88 | ## Introduction |
| 89 | |
| 90 | We are excited to open-source **Qwen-Image-2.1**, a unified text-to-image generation and image editing model in the Qwen family. With just **7B parameters in its visual generation component** (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility. |
| 91 | |
| 92 | Four key improvements define this release: |
| 93 | |
| 94 | - **Compact and Efficient**: a lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost. |
| 95 | - **Native Transparency, Unified Creation and Editing**: generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs, all in one model. |
| 96 | - **Versatile Editing**: support up to **10 reference images**, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products. |
| 97 | - **Realistic Textures and Refined Aesthetics**: improved typography, portrait lighting, and fine details for more visually compelling results. |
| 98 | |
| 99 | <p align="center"> |
| 100 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-01.png" width="100%"/> |
| 101 | </p> |
| 102 | |
| 103 | For more details, see the [GitHub repo](https://github.com/QwenLM/Qwen-Image-2.1) and [Blog](https://qwen.ai/blog?id=qwen-image-2.1). |
| 104 | |
| 105 | ## Quick Start |
| 106 | |
| 107 | ### Installation |
| 108 | |
| 109 | ```bash |
| 110 | pip install torch>=2.4.0 |
| 111 | pip install transformers>=5.17 |
| 112 | pip install git+https://github.com/huggingface/diffusers |
| 113 | pip install accelerate pillow |
| 114 | ``` |
| 115 | |
| 116 | ### Text-to-Image |
| 117 | |
| 118 | ```python |
| 119 | import torch |
| 120 | from diffusers import QwenImage21Pipeline |
| 121 | |
| 122 | pipe = QwenImage21Pipeline.from_pretrained( |
| 123 | "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16 |
| 124 | ).to("cuda") |
| 125 | |
| 126 | image = pipe( |
| 127 | prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement", |
| 128 | width=2048, height=2048, |
| 129 | num_inference_steps=40, |
| 130 | generator=torch.Generator("cuda").manual_seed(42), |
| 131 | ).images[0] |
| 132 | |
| 133 | image.save("t2i_example.png") |
| 134 | ``` |
| 135 | |
| 136 | ### Image Editing |
| 137 | |
| 138 | ```python |
| 139 | import torch |
| 140 | from PIL import Image |
| 141 | from diffusers import QwenImage21Pipeline |
| 142 | |
| 143 | pipe = QwenImage21Pipeline.from_pretrained( |
| 144 | "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16 |
| 145 | ).to("cuda") |
| 146 | |
| 147 | input_image = Image.open("input.png") |
| 148 | |
| 149 | image = pipe( |
| 150 | prompt="Change the background to a sunset beach", |
| 151 | image=input_image, |
| 152 | num_inference_steps=40, |
| 153 | generator=torch.Generator("cuda").manual_seed(42), |
| 154 | ).images[0] |
| 155 | |
| 156 | image.save("edit_example.png") |
| 157 | ``` |
| 158 | |
| 159 | ### Transparent Image Generation (RGBA) |
| 160 | |
| 161 | Use the recommended prompt format for transparent images: |
| 162 | |
| 163 | ```python |
| 164 | image = pipe( |
| 165 | prompt="This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.", |
| 166 | width=2048, height=2048, |
| 167 | num_inference_steps=40, |
| 168 | generator=torch.Generator("cuda").manual_seed(42), |
| 169 | ).images[0] |
| 170 | |
| 171 | image.save("transparent_example.png") |
| 172 | ``` |
| 173 | |
| 174 | ### Supported Aspect Ratios |
| 175 | |
| 176 | ```python |
| 177 | aspect_ratios = { |
| 178 | "1:1": (2048, 2048), |
| 179 | "4:3": (2400, 1792), |
| 180 | "3:4": (1792, 2400), |
| 181 | "3:2": (2528, 1696), |
| 182 | "2:3": (1696, 2528), |
| 183 | "16:9": (2752, 1536), |
| 184 | "9:16": (1536, 2752), |
| 185 | } |
| 186 | ``` |
| 187 | |
| 188 | ### Memory Optimization |
| 189 | |
| 190 | ```python |
| 191 | pipe = QwenImage21Pipeline.from_pretrained( |
| 192 | "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16 |
| 193 | ) |
| 194 | pipe.enable_model_cpu_offload() |
| 195 | ``` |
| 196 | |
| 197 | ## Showcase |
| 198 | |
| 199 | <p align="center"> |
| 200 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-04.png" width="30%"/> |
| 201 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-05.png" width="30%"/> |
| 202 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-06.png" width="30%"/> |
| 203 | </p> |
| 204 | <p align="center"><em>Native transparent image generation</em></p> |
| 205 | |
| 206 | <p align="center"> |
| 207 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-15.png" width="100%"/> |
| 208 | </p> |
| 209 | <p align="center"><em>Group photograph generated from six portrait references</em></p> |
| 210 | |
| 211 | <p align="center"> |
| 212 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-43.png" width="48%"/> |
| 213 | <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-44.png" width="48%"/> |
| 214 | </p> |
| 215 | <p align="center"><em>Text rendering</em></p> |
| 216 | |
| 217 | ## License |
| 218 | |
| 219 | This model is licensed under the [Qwen Research License Agreement](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE). |
| 220 | |