Instructions to use ThakiCloud/Qwen-Image-2.1-FewStep-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ThakiCloud/Qwen-Image-2.1-FewStep-v0.1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("ThakiCloud/Qwen-Image-2.1-FewStep-v0.1") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen-Image-2.1 FewStep v0.1
8-step and 5-step LoRA adapters for Qwen-Image-2.1 that score higher than the 40-step base model on PickScore, and beat PrunaAI's 5/8-step adapters on every prompt set we measured.
Two LoRA adapters (rank 64) distilled from Qwen/Qwen-Image-2.1 with backward-simulated DMD2
(distribution matching, two-timescale fake-score updates, no GAN term). They load on top of the
unchanged diffusers pipeline and run without CFG. Improved using Qwen.
🇰🇷 한국어 카드: MODEL_CARD.ko.md
| Base | Qwen/Qwen-Image-2.1 (7B single-stream DiT + Qwen3-VL-8B text encoder) |
| Method | DMD2 with backward simulation on the deployment sigma schedule · LoRA r=64 on attn.to_q/k/v/to_out.0 + img_mlp.gate_layer/proj/out (all 32 blocks) · 1,200 generator updates, global batch 4, 1024² |
| Training prompts | 33k: detailed FLUX-style prompts (MIT) + short SD-style prompts + synthetic text-rendering prompts (en/ko/zh, 8% of draws for the 8-step adapter). Targets are the base model itself; no external images |
| Files | qif_qwen_image_2.1_8step_v0.1.safetensors · qif_qwen_image_2.1_5step_v0.1.safetensors (335 MB each, same key layout as other Qwen-Image-2.1 LoRAs) |
| License | Qwen Research License (non-commercial) — inherited from the base |
⚠️ Text-to-image only. Qwen-Image-2.1 also edits images with up to 10 references; these adapters were trained and evaluated on T2I only. Editing quality with the adapter loaded is not verified.
⚠️ The OCR numbers are partly in-distribution. The 100 text-rendering eval prompts are held-out (carrier, style, word) combinations, but they share the templates and word pool with the training prompts. An out-of-distribution text benchmark (LongText-Bench) is the next measurement.
0. How to use it
import torch
from diffusers import QwenImage21Pipeline, FlowMatchEulerDiscreteScheduler
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16).to("cuda")
pipe.load_lora_weights("ThakiCloud/Qwen-Image-2.1-FewStep-v0.1",
weight_name="qif_qwen_image_2.1_8step_v0.1.safetensors")
# the adapter has its own sigma list: turn the base scheduler's dynamic shift off so it is not applied twice
pipe.scheduler = FlowMatchEulerDiscreteScheduler.from_config(
pipe.scheduler.config, use_dynamic_shifting=False, shift=1.0, shift_terminal=None)
SIGMAS = {8: [1.0, 14/15, 6/7, 10/13, 2/3, 6/11, 0.4, 2/9],
5: [1.0, 0.94, 6/7, 2/3, 0.4]}
image = pipe('A neon shop sign that reads "OPEN LATE", rainy night',
height=1024, width=1024, num_inference_steps=8, sigmas=SIGMAS[8]).images[0]
For the 5-step adapter use weight_name="qif_qwen_image_2.1_5step_v0.1.safetensors",
num_inference_steps=5, sigmas=SIGMAS[5]. No negative prompt, true_cfg_scale stays at 1.
For the fastest path, fuse the LoRA (pipe.fuse_lora()) and compile the transformer (§2).
The 8-step sigma list is the base scheduler's own shift at 1024² (exp(μ)=2.0) applied to a uniform grid, so the adapter matches what the base model expects at that resolution. The adapters were trained at 1024²; other resolutions are untested.
1. Quality — 500 prompts, 1024², same seed per prompt
PartiPrompts (200, stratified) · DrawBench (200) · text rendering (100, en/ko/zh). PickScore and CLIP-H are means over all 500 images; OCR is on the text-rendering slice, transcribed by Qwen3-VL-8B and counted as a hit when the requested string appears in the transcription.
| steps | PickScore | CLIP-H | OCR exact | OCR CER | PickScore win vs Pruna (same steps) | |
|---|---|---|---|---|---|---|
| Qwen-Image-2.1 (official recipe) | 40 | 22.31 | 35.02 | 95% | 3.2% | — |
| Qwen-Image-2.1, no adapter | 8 | 21.47 | 34.24 | 81% | 9.6% | 17% |
| PrunaAI/Pruna-Qwen-Image-2.1 | 8 | 22.02 | 35.00 | 91% | 4.4% | — |
| this repo, 8-step | 8 | 22.47 | 34.66 | 99% | 0.5% | 68% |
| Qwen-Image-2.1, no adapter | 5 | 21.35 | 34.29 | 82% | 10.2% | 14% |
| PrunaAI/Pruna-Qwen-Image-2.1 | 5 | 21.94 | 35.10 | 87% | 6.2% | — |
| this repo, 5-step | 5 | 22.34 | 34.35 | 99% | 0.1% | 69% |
- 8-step vs Pruna 8-step: PickScore +0.46 (95% bootstrap CI +0.38 … +0.53).
- 8-step vs the 40-step base: PickScore +0.16 (CI +0.09 … +0.23), per-prompt win rate 59%.
- Where it is weaker: CLIP-H (prompt–image alignment) is 0.3–0.8 below Pruna and the base. On PartiPrompts the 8-step adapter is tied with Pruna (32.95 vs 32.93); the gap comes from DrawBench and the text slice. PickScore rewards the higher contrast DMD tends to produce, so read the two columns together.
2. Speed — one NVIDIA B200, 1024², batch 1, bf16
Wall time per image including text encoding and VAE decoding, median of 20 after warm-up.
| s / image | vs 40-step | |
|---|---|---|
| Qwen-Image-2.1, 40 steps (KV cache on) | 3.69 | 1.0× |
| 8-step, LoRA unfused | 0.92 | 4.0× |
| 8-step, LoRA fused | 0.82 | 4.5× |
8-step, LoRA fused + torch.compile |
0.66 | 5.5× |
| 5-step, LoRA unfused | 0.62 | 5.9× |
| 5-step, LoRA fused | 0.56 | 6.6× |
5-step, LoRA fused + torch.compile |
0.45 | 8.2× |
Unfused, the adapters run at the same speed as Pruna's (0.94 s / 0.64 s in the same harness); the extra speed comes from fusing and compiling, which any LoRA can use.
3. How it was trained
One frozen copy of the base carries two LoRA branches: the student (shipped) and a fake score that tracks the student's output distribution; the base without LoRA is the real score. Each generator update rolls the student out on the deployment sigma schedule for a random number of steps without gradient, takes one step with gradient, re-noises the prediction at a noise level drawn with the base's shift, and moves it along the difference between the fake and real denoised estimates. The fake branch is updated three times per generator update. One B200 fits the whole setup in 42 GB; the 8-step adapter took 2.7 hours on one GPU.
4. Limitations
- T2I at 1024² only; editing and the base's native 2048² are not evaluated.
- Text-rendering gains are measured on templates close to the training prompts (see the note above).
- Automatic metrics only (PickScore, CLIP-H, VLM OCR); no human preference study yet.
- Non-commercial license, inherited from Qwen-Image-2.1.
License
Qwen Research License Agreement — research and evaluation only. See LICENSE and NOTICE.
Commercial use needs a separate license from the Qwen team. Built on Qwen-Image-2.1 by
Alibaba/Tongyi Lab; the adapters are a modification of it.
- Downloads last month
- 494
Model tree for ThakiCloud/Qwen-Image-2.1-FewStep-v0.1
Base model
Qwen/Qwen-Image-2.1