YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
HunyuanImage-3.0-Instruct-Distil · MXFP8 Model
This is an MXFP8 quantized build of tencent/HunyuanImage-3.0-Instruct-Distil, exported by
AutoRound with --format auto_round (the vLLM INC path).
Both stages of the model — the AR (autoregressive language model) and the DiT (diffusion transformer) —
have been verified end-to-end on vLLM + vLLM-Omni.
Generated output, not a ground-truth reference: AR+DiT full pipeline,
seed=42, 8 inference steps, promptA cute cat(the AR stage first emits the CoT and the image-ratio tokens, whose KV cache is then handed to the DiT stage).
Overview
| Field | Value |
|---|---|
| Base model | tencent/HunyuanImage-3.0-Instruct-Distil (the Distil build: cfg_distilled=true, use_meanflow=true) |
| Quantization scheme | AutoRound MXFP8 (--scheme MXFP8): 32-element blocks, 8-bit, symmetric, shared exponent scale (E8M0) |
| Export format | --format auto_round → quantization_config.quant_method = "auto-round", data_type = "mx_fp" (vLLM INC path) |
| Quantization tool | auto-round 0.15.0 (--model_free, no calibration data required) |
| Disk size | 86 GB (BF16 base model: 158 GB, ≈ 0.54×) |
| Weight dtype | FP8-E4M3 values stored as BF16 + uint8 MX block scales |
| Quantized | Linear layers of the AR and the DiT (including the MoE expert gate_and_up_proj / down_proj), scope block_name_to_quantize = "model.layers" |
| Kept in BF16 | vision (ViT), wte / lm_head, guidance_emb, timestep_emb, timestep_r_emb, final_layer, and every layer marked bits=16 in extra_config — including the MoE router of all 32 blocks (…mlp.gate.wg) |
Key fields of quantization_config (see config.json):
{
"quant_method": "auto-round",
"data_type": "mx_fp",
"bits": 8,
"group_size": 32,
"sym": true,
"packing_format": "auto_round:llm_compressor",
"block_name_to_quantize": "model.layers",
"extra_config": { ...220 entries, 32 of them "model.layers.N.mlp.gate.wg": {"bits": 16, ...} ... }
}
Quantization command
# Requires auto-round >= 0.15.0 (this build used 0.15.0) and a BF16 copy of the base model.
auto-round \
--model_name tencent/HunyuanImage-3.0-Instruct-Distil \
--model_free \
--scheme MXFP8 \
--ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte" \
--format auto_round \
--device cuda:0 \
--output_dir ./HunyuanImage-3.0-Instruct-Distil-MXFP8
--ignore_layersmust excludevision: the ViT'smlp.fc2has an input dimension of 4304, which is not a multiple of the 32-element MXFP8 block size (MXFP8 requires input_size_per_partition (2152) to be divisible by 32).- AutoRound records the MoE router in
extra_config("model.layers.N.mlp.gate.wg": {"bits": 16}) for all 32 blocks.
Inference environment
| Component | Version |
|---|---|
| vLLM | 0.29.0 (the auto_round / INC path requires ≥ 0.29.0, see above) |
| vLLM-Omni | latest main (this build was validated at 1c7476ec, version string 0.29.0rc2.dev161+g1c7476ec1, installed editable) |
| PyTorch | 2.13.0+cu132 |
| FlashInfer | 0.6.18 |
# Install (same as the environment this build was validated in)
pip install vllm==0.29.0
pip install -e /path/to/vllm-omni # main branch
# Sanity check: make sure the expected vllm-omni is the one being imported
python -c "import vllm, vllm_omni, os; print(vllm.__version__); print(os.path.dirname(vllm_omni.__file__))"
Environment variables and working directory
Inference example: AR + DiT full pipeline (text2img)
Use vLLM-Omni's offline entry point examples/offline_inference/text_to_image/text_to_image.py with a
two-stage deploy YAML (AR = stage 0, DiT = stage 1, KV handed over through shared memory).
0. Prepare the deploy YAML (save the block below as hunyuan_image_3_moe.yaml)
# AR (stage 0) + DiT (stage 1), 4 GPUs: 2 for AR, 2 for DiT
pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true
connectors:
shared_memory_connector:
name: SharedMemoryConnector
stages:
- stage_id: 0
is_comprehension: true
final_output: true
final_output_type: text
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
max_num_batched_tokens: 32768
devices: "0,1"
tensor_parallel_size: 2
hf_overrides:
rope_parameters:
mrope_section: [0, 32, 32]
rope_type: default
omni_kv_config:
need_send_cache: true
output_connectors:
to_stage_1: shared_memory_connector
default_sampling_params:
temperature: 0.0
top_p: 1
top_k: -1
max_tokens: 8192
detokenize: true
skip_special_tokens: false
include_stop_str_in_output: true
- stage_id: 1
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
devices: "2,3"
distributed_executor_backend: "mp"
omni_kv_config:
need_recv_cache: true
parallel_config:
tensor_parallel_size: 2
enable_expert_parallel: true
input_connectors:
from_stage_0: shared_memory_connector
default_sampling_params:
num_inference_steps: 8
guidance_scale: 0
edges:
- from: 0
to: 1
window_size: -1
max_inflight: 1
devicesare local indices and are meant to be combined withCUDA_VISIBLE_DEVICES.- With only 2 GPUs, set both stages to
tensor_parallel_size: 1withdevices: "0"/"1"(at TP=1 each stage holds the whole stage's weights, ~85 GiB — see "Memory footprint").shared_memory_connectorreplaces the official YAML's RDMA/Mooncake transport for single-host runs, so no external dependency is needed.guidance_scale: 0on the DiT stage is only a fallback; the value actually used comes from the CLI.
Memory footprint
From the AR+DiT run logs (Model loading took …):
| Stage | Per card at TP=2 | Process total per card |
|---|---|---|
| stage 0 = AR | 42.7 GiB | ~48 GiB |
| stage 1 = DiT | 43.4 GiB | ~45 GiB |
The whole checkpoint is 84.98 GiB as reported by vLLM (86 GB per du -sh). Therefore:
- 4 GPUs (AR TP2 + DiT TP2): ~43–49 GiB per card, plenty of headroom.
- 2 GPUs (AR TP1 + DiT TP1): ~85 GiB per card, fits on a 141 GB card (verified).
- The BF16 base model cannot run on 2 GPUs: its AR stage already fills a card at TP=2, so the base model needs at least 4 GPUs (AR TP2 + DiT TP2).
1. Run
export CUDA_VISIBLE_DEVICES=0,1,2,3 # four free GPUs
cd /tmp # cwd must be neutral
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /path/to/HunyuanImage-3.0-Instruct-Distil-MXFP8 \
--deploy-config ./hunyuan_image_3_moe.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 \
--guidance-scale 5.0 \
--output ./output.png
On success the log ends with Saved generated image to ./output.png and contains
Using 'MARLIN' MxFp8 MoE backend.
2. Parameter notes (important)
--num-inference-steps 8: this is an 8-step distilled model. Do not use the 50 steps of the Instruct build.--guidance-scaleis a real input for this model. The checkpoint declarescfg_distilled=true, and vLLM-Omni feeds1000 × guidance_scaleinto the DiT as a guidance embedding — this value materially changes the output. The verification image here used5.0. If you want to compare against the reference image shipped with vLLM-Omni (tests/assets/hunyuan_image3/hunyuan_image_distill_ref.png), use 2.5 (that is what the Distil branch oftests/e2e/accuracy/test_hunyuan_image3.pyuses).--promptonly reaches the AR stage; the AR emits the CoT and the image-ratio tokens first, then hands its KV cache to the DiT.
Results
Same machine, same prompt (A cute cat), seed=42, 8 steps, guidance-scale 5.0, TP=2 for
both stages, compared against the BF16 base model on the same AR+DiT pipeline:
- Downloads last month
- 20

