YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

HunyuanImage-3.0-Instruct-Distil · MXFP8 Model

This is an MXFP8 quantized build of tencent/HunyuanImage-3.0-Instruct-Distil, exported by AutoRound with --format auto_round (the vLLM INC path). Both stages of the model — the AR (autoregressive language model) and the DiT (diffusion transformer) — have been verified end-to-end on vLLM + vLLM-Omni.

AR+DiT full-pipeline output

Generated output, not a ground-truth reference: AR+DiT full pipeline, seed=42, 8 inference steps, prompt A cute cat (the AR stage first emits the CoT and the image-ratio tokens, whose KV cache is then handed to the DiT stage).

Overview

Field Value
Base model tencent/HunyuanImage-3.0-Instruct-Distil (the Distil build: cfg_distilled=true, use_meanflow=true)
Quantization scheme AutoRound MXFP8 (--scheme MXFP8): 32-element blocks, 8-bit, symmetric, shared exponent scale (E8M0)
Export format --format auto_round → quantization_config.quant_method = "auto-round", data_type = "mx_fp" (vLLM INC path)
Quantization tool auto-round 0.15.0 (--model_free, no calibration data required)
Disk size 86 GB (BF16 base model: 158 GB, ≈ 0.54×)
Weight dtype FP8-E4M3 values stored as BF16 + uint8 MX block scales
Quantized Linear layers of the AR and the DiT (including the MoE expert gate_and_up_proj / down_proj), scope block_name_to_quantize = "model.layers"
Kept in BF16 vision (ViT), wte / lm_head, guidance_emb, timestep_emb, timestep_r_emb, final_layer, and every layer marked bits=16 in extra_config — including the MoE router of all 32 blocks (…mlp.gate.wg)

Key fields of quantization_config (see config.json):

{
  "quant_method": "auto-round",
  "data_type": "mx_fp",
  "bits": 8,
  "group_size": 32,
  "sym": true,
  "packing_format": "auto_round:llm_compressor",
  "block_name_to_quantize": "model.layers",
  "extra_config": { ...220 entries, 32 of them "model.layers.N.mlp.gate.wg": {"bits": 16, ...} ... }
}

Quantization command

# Requires auto-round >= 0.15.0 (this build used 0.15.0) and a BF16 copy of the base model.
auto-round \
  --model_name tencent/HunyuanImage-3.0-Instruct-Distil \
  --model_free \
  --scheme MXFP8 \
  --ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte" \
  --format auto_round \
  --device cuda:0 \
  --output_dir ./HunyuanImage-3.0-Instruct-Distil-MXFP8
  • --ignore_layers must exclude vision: the ViT's mlp.fc2 has an input dimension of 4304, which is not a multiple of the 32-element MXFP8 block size (MXFP8 requires input_size_per_partition (2152) to be divisible by 32).
  • AutoRound records the MoE router in extra_config ("model.layers.N.mlp.gate.wg": {"bits": 16}) for all 32 blocks.

Inference environment

Component Version
vLLM 0.29.0 (the auto_round / INC path requires ≥ 0.29.0, see above)
vLLM-Omni latest main (this build was validated at 1c7476ec, version string 0.29.0rc2.dev161+g1c7476ec1, installed editable)
PyTorch 2.13.0+cu132
FlashInfer 0.6.18
# Install (same as the environment this build was validated in)
pip install vllm==0.29.0
pip install -e /path/to/vllm-omni        # main branch

# Sanity check: make sure the expected vllm-omni is the one being imported
python -c "import vllm, vllm_omni, os; print(vllm.__version__); print(os.path.dirname(vllm_omni.__file__))"

Environment variables and working directory

Inference example: AR + DiT full pipeline (text2img)

Use vLLM-Omni's offline entry point examples/offline_inference/text_to_image/text_to_image.py with a two-stage deploy YAML (AR = stage 0, DiT = stage 1, KV handed over through shared memory).

0. Prepare the deploy YAML (save the block below as hunyuan_image_3_moe.yaml)

# AR (stage 0) + DiT (stage 1), 4 GPUs: 2 for AR, 2 for DiT
pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true

connectors:
  shared_memory_connector:
    name: SharedMemoryConnector

stages:
  - stage_id: 0
    is_comprehension: true
    final_output: true
    final_output_type: text
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    max_num_batched_tokens: 32768
    devices: "0,1"
    tensor_parallel_size: 2
    hf_overrides:
      rope_parameters:
        mrope_section: [0, 32, 32]
        rope_type: default
    omni_kv_config:
      need_send_cache: true
    output_connectors:
      to_stage_1: shared_memory_connector
    default_sampling_params:
      temperature: 0.0
      top_p: 1
      top_k: -1
      max_tokens: 8192
      detokenize: true
      skip_special_tokens: false
      include_stop_str_in_output: true

  - stage_id: 1
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    devices: "2,3"
    distributed_executor_backend: "mp"
    omni_kv_config:
      need_recv_cache: true
    parallel_config:
      tensor_parallel_size: 2
      enable_expert_parallel: true
    input_connectors:
      from_stage_0: shared_memory_connector
    default_sampling_params:
      num_inference_steps: 8
      guidance_scale: 0

edges:
  - from: 0
    to: 1
    window_size: -1
    max_inflight: 1
  • devices are local indices and are meant to be combined with CUDA_VISIBLE_DEVICES.
  • With only 2 GPUs, set both stages to tensor_parallel_size: 1 with devices: "0" / "1" (at TP=1 each stage holds the whole stage's weights, ~85 GiB — see "Memory footprint").
  • shared_memory_connector replaces the official YAML's RDMA/Mooncake transport for single-host runs, so no external dependency is needed.
  • guidance_scale: 0 on the DiT stage is only a fallback; the value actually used comes from the CLI.

Memory footprint

From the AR+DiT run logs (Model loading took …):

Stage Per card at TP=2 Process total per card
stage 0 = AR 42.7 GiB ~48 GiB
stage 1 = DiT 43.4 GiB ~45 GiB

The whole checkpoint is 84.98 GiB as reported by vLLM (86 GB per du -sh). Therefore:

  • 4 GPUs (AR TP2 + DiT TP2): ~43–49 GiB per card, plenty of headroom.
  • 2 GPUs (AR TP1 + DiT TP1): ~85 GiB per card, fits on a 141 GB card (verified).
  • The BF16 base model cannot run on 2 GPUs: its AR stage already fills a card at TP=2, so the base model needs at least 4 GPUs (AR TP2 + DiT TP2).

1. Run

export CUDA_VISIBLE_DEVICES=0,1,2,3      # four free GPUs

cd /tmp                                   # cwd must be neutral
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model            /path/to/HunyuanImage-3.0-Instruct-Distil-MXFP8 \
  --deploy-config    ./hunyuan_image_3_moe.yaml \
  --prompt           "A cute cat" \
  --num-inference-steps 8 \
  --guidance-scale   5.0 \
  --output           ./output.png

On success the log ends with Saved generated image to ./output.png and contains Using 'MARLIN' MxFp8 MoE backend.

2. Parameter notes (important)

  • --num-inference-steps 8: this is an 8-step distilled model. Do not use the 50 steps of the Instruct build.
  • --guidance-scale is a real input for this model. The checkpoint declares cfg_distilled=true, and vLLM-Omni feeds 1000 × guidance_scale into the DiT as a guidance embedding — this value materially changes the output. The verification image here used 5.0. If you want to compare against the reference image shipped with vLLM-Omni (tests/assets/hunyuan_image3/hunyuan_image_distill_ref.png), use 2.5 (that is what the Distil branch of tests/e2e/accuracy/test_hunyuan_image3.py uses).
  • --prompt only reaches the AR stage; the AR emits the CoT and the image-ratio tokens first, then hands its KV cache to the DiT.

Results

Same machine, same prompt (A cute cat), seed=42, 8 steps, guidance-scale 5.0, TP=2 for both stages, compared against the BF16 base model on the same AR+DiT pipeline:

Output
This build (MXFP8) mxfp8
BF16 base model (reference) bf16
Downloads last month
20
Safetensors
Model size
83B params
Tensor type
BF16
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support