README.md
8.8 KB · 220 lines · markdown Raw
1 ---
2 base_model: Qwen/Qwen-Image-2.1
3 base_model_relation: quantized
4 license: other
5 license_name: qwen-research
6 license_link: https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE
7 language:
8 - en
9 - zh
10 pipeline_tag: text-to-image
11 tags:
12 - gguf
13 - quantized
14 - unsloth
15 - qwen
16 - image-generation
17 widget:
18 - text: Photorealistic editorial photograph of a woman barista making a latte in a modern, minimalist café on a sunny tropical morning.
19 output:
20 url: assets/cafe.png
21 - text: A lone astronaut crossing a dark, frozen lake beneath enormous rings stretching across an alien sky. Fine cracks visible under the translucent ice, distant mountains, soft blue twilight, cinematic scale, photorealistic detail.
22 output:
23 url: assets/spaces.png
24 ---
25 # Read our How to [Run Qwen-Image-2.1 Guide!](https://unsloth.ai/docs/models/qwen-image-2.1) 💜
26
27 This is a GGUF quantized version of [Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1). <br>
28 unsloth/Qwen-Image-2.1-GGUF uses [Unsloth Dynamic 2.0](https://docs.unsloth.ai/basics/unsloth-dynamic-2.0-ggufs) methodology for SOTA performance.
29
30 - Important layers are upcasted to higher precision, per tensor, from a measured sensitivity scan.
31 - Run these with [Unsloth Desktop](https://github.com/unslothai/unsloth), stable-diffusion.cpp and more. A GGUF is the denoiser only, so it needs the VAE and the Qwen3-VL text encoder alongside it.
32 - VAE: [unsloth/Qwen-Image-2.1-FP8](https://huggingface.co/unsloth/Qwen-Image-2.1-FP8) `vae/qwen_image_2.1_vae_bf16.safetensors`. Text encoder: [unsloth/Qwen3-VL-8B-Instruct-GGUF](https://huggingface.co/unsloth/Qwen3-VL-8B-Instruct-GGUF) `Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf`, the Dynamic 2.0 4-bit rung rather than the uniform `Q4_K_M`. Measured against the `Q4_K_M` encoder at a shared seed, with the denoiser and VAE held fixed: LPIPS 0.029, SSIM 0.959, 5.15 GB vs 5.03 GB, 36.5 s vs 39.0 s.
33 <div>
34 <div style="display: flex; gap: 5px; align-items: center; ">
35 <a href="https://github.com/unslothai/unsloth/">
36 <img src="https://github.com/unslothai/unsloth/raw/main/images/unsloth%20new%20logo.png" width="133">
37 </a>
38 <a href="https://discord.gg/unsloth">
39 <img src="https://github.com/unslothai/unsloth/raw/main/images/Discord%20button.png" width="173">
40 </a>
41 <a href="https://unsloth.ai/docs/models/qwen-image-2.1">
42 <img src="https://raw.githubusercontent.com/unslothai/unsloth/refs/heads/main/images/documentation%20green%20button.png" width="143">
43 </a>
44 </div>
45 </div>
46 See below for image editing operating inside of Unsloth Desktop:
47 <img width="600" alt="qwen-image-2.1 unsloth desktop" src="https://unsloth.ai/docs/~gitbook/image?url=https%3A%2F%2F3215535692-files.gitbook.io%2F%7E%2Ffiles%2Fv0%2Fb%2Fgitbook-x-prod.appspot.com%2Fo%2Fspaces%252FxhOjnexMCB3dmuQFQ2Zq%252Fuploads%252F8E4zaQZmeYBhYIXkvveM%252F01-2edit-multi-image.webp%3Falt%3Dmedia%26token%3Dad7d102c-b304-4c75-b66d-4a68cd518a37&width=768&dpr=3&quality=100&sign=a5a37f593c66ddc27c847bab0d138eda&sv=3" />
48
49
50 ```bash
51 sd-cli --diffusion-model qwen-image-2.1-Q4_K_M.gguf \
52 --vae qwen_image_2.1_vae_bf16.safetensors \
53 --llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf \
54 -p "a cartoon sloth mascot waving, flat vector illustration, bright colours" \
55 --steps 20 --cfg-scale 6.0 --sampling-method euler -W 1024 -H 1024 --diffusion-fa \
56 -o out.png
57 ```
58
59 ### Samples
60
61 Rendered with the Q4_K_M denoiser and the Q4_K_M text encoder, 1024x1024, 20 steps, cfg 6.0, euler.
62
63 <table>
64 <tr>
65 <td><img src="assets/cafe.png" width="200"></td>
66 <td><img src="assets/spaces.png" width="200"></td>
67 </tr>
68 <tr>
69 <td><img src="assets/cinema.png" width="200"></td>
70 <td><img src="assets/cardesert.png" width="200"></td>
71 </tr>
72 </table>
73
74 ---
75
76 <p align="center">
77 <img src="https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-Image/image2.1/logo.png" width="400"/>
78 </p>
79 <p align="center">
80 🤖 <a href="https://modelscope.cn/models/Qwen/Qwen-Image-2.1">ModelScope</a>&nbsp;&nbsp;|
81 &nbsp;&nbsp;🤗 <a href="https://huggingface.co/Qwen/Qwen-Image-2.1">HuggingFace</a>&nbsp;&nbsp;|
82 &nbsp;&nbsp;📑 <a href="https://qwen.ai/blog?id=qwen-image-2.1">Blog</a>&nbsp;&nbsp;|
83 &nbsp;&nbsp;🖥️ <a href="https://huggingface.co/spaces/Qwen/Qwen-Image-2.1">Demo</a>&nbsp;&nbsp;|
84 &nbsp;&nbsp;🫨 <a href="https://discord.gg/BEYSk3pkSu">Discord</a>&nbsp;&nbsp;|
85 &nbsp;&nbsp;💬 <a href="https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/assets/qr.png">WeChat</a>
86 </p>
87
88 ## Introduction
89
90 We are excited to open-source **Qwen-Image-2.1**, a unified text-to-image generation and image editing model in the Qwen family. With just **7B parameters in its visual generation component** (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
91
92 Four key improvements define this release:
93
94 - **Compact and Efficient**: a lightweight architecture with mixed-granularity attention and prefix KV cache reuse delivers strong image quality at low computational cost.
95 - **Native Transparency, Unified Creation and Editing**: generate regular or transparent (RGBA) images from text, edit transparent layers, and extract subjects from photographs, all in one model.
96 - **Versatile Editing**: support up to **10 reference images**, specify local edits via circles, painted annotations, or separate masks, and preserve identity for people and products.
97 - **Realistic Textures and Refined Aesthetics**: improved typography, portrait lighting, and fine details for more visually compelling results.
98
99 <p align="center">
100 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-01.png" width="100%"/>
101 </p>
102
103 For more details, see the [GitHub repo](https://github.com/QwenLM/Qwen-Image-2.1) and [Blog](https://qwen.ai/blog?id=qwen-image-2.1).
104
105 ## Quick Start
106
107 ### Installation
108
109 ```bash
110 pip install torch>=2.4.0
111 pip install transformers>=5.17
112 pip install git+https://github.com/huggingface/diffusers
113 pip install accelerate pillow
114 ```
115
116 ### Text-to-Image
117
118 ```python
119 import torch
120 from diffusers import QwenImage21Pipeline
121
122 pipe = QwenImage21Pipeline.from_pretrained(
123 "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
124 ).to("cuda")
125
126 image = pipe(
127 prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
128 width=2048, height=2048,
129 num_inference_steps=40,
130 generator=torch.Generator("cuda").manual_seed(42),
131 ).images[0]
132
133 image.save("t2i_example.png")
134 ```
135
136 ### Image Editing
137
138 ```python
139 import torch
140 from PIL import Image
141 from diffusers import QwenImage21Pipeline
142
143 pipe = QwenImage21Pipeline.from_pretrained(
144 "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
145 ).to("cuda")
146
147 input_image = Image.open("input.png")
148
149 image = pipe(
150 prompt="Change the background to a sunset beach",
151 image=input_image,
152 num_inference_steps=40,
153 generator=torch.Generator("cuda").manual_seed(42),
154 ).images[0]
155
156 image.save("edit_example.png")
157 ```
158
159 ### Transparent Image Generation (RGBA)
160
161 Use the recommended prompt format for transparent images:
162
163 ```python
164 image = pipe(
165 prompt="This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.",
166 width=2048, height=2048,
167 num_inference_steps=40,
168 generator=torch.Generator("cuda").manual_seed(42),
169 ).images[0]
170
171 image.save("transparent_example.png")
172 ```
173
174 ### Supported Aspect Ratios
175
176 ```python
177 aspect_ratios = {
178 "1:1": (2048, 2048),
179 "4:3": (2400, 1792),
180 "3:4": (1792, 2400),
181 "3:2": (2528, 1696),
182 "2:3": (1696, 2528),
183 "16:9": (2752, 1536),
184 "9:16": (1536, 2752),
185 }
186 ```
187
188 ### Memory Optimization
189
190 ```python
191 pipe = QwenImage21Pipeline.from_pretrained(
192 "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
193 )
194 pipe.enable_model_cpu_offload()
195 ```
196
197 ## Showcase
198
199 <p align="center">
200 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-04.png" width="30%"/>
201 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-05.png" width="30%"/>
202 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-06.png" width="30%"/>
203 </p>
204 <p align="center"><em>Native transparent image generation</em></p>
205
206 <p align="center">
207 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-15.png" width="100%"/>
208 </p>
209 <p align="center"><em>Group photograph generated from six portrait references</em></p>
210
211 <p align="center">
212 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-43.png" width="48%"/>
213 <img src="https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen-Image/image2.1/images/example-44.png" width="48%"/>
214 </p>
215 <p align="center"><em>Text rendering</em></p>
216
217 ## License
218
219 This model is licensed under the [Qwen Research License Agreement](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE).
220