YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

HunyuanImage-3.0-Instruct-Distil · MXFP4 mixed-precision with attention kept in BF16 (auto-round / vLLM INC path)

Same MXFP4-for-routed-experts recipe as -MXFP4-ar, plus one deliberate change: self_attn.qkv_proj and self_attn.o_proj are added to --ignore_layers, so attention stays BF16.

Recipe summary: routed experts → MXFP4 · mlp.shared_mlp → MXFP8 · self_attn → BF16 · everything else (ViT, VAE, MoE router, embeddings, DiT heads) → BF16/FP32.

AR+DiT full-pipeline output

Generated output, not a ground-truth reference. AR+DiT on 2 GPUs (AR TP1 + DiT TP1), prompt A cute cat, seed=42, 8 steps, guidance_scale=5.0. Captured 2026-09-24 08:54, exit=0.


⚠️ Read this first

1. This build is -MXFP4-ar + "un-quantize attention" — and it buys you almost nothing

Measured in one session, same topology, DiT-only TP2 (so attention precision is the only variable):

pair PSNR SSIM@1 SSIM@1/8
this build ↔ -MXFP4-ar 32.97 0.9802 0.9806
this build ↔ BF16 19.60 0.8494 0.6699
-MXFP4-ar ↔ BF16 19.51 0.8421 0.6627
-MXFP8 (all 8-bit) ↔ BF16 32.77 0.9760 0.9813

⇒ Restoring BF16 attention costs +1.30 GB on disk and +≈1.0 GiB/GPU at load time, and changes the picture at the level of run-to-run numerical noise. The composition divergence of MXFP4 comes from the experts' 4-bit codebook, not from attention precision. Keep this build only if you have a specific reason to protect attention (e.g. a downstream experiment), not as a fidelity fix.

2. A stock --format auto_round export of this recipe does NOT load — the metadata fix is required

AutoRound 0.15.0 writes the --layer_config pattern into extra_config verbatim (key = ".mlp.experts."), which vLLM's INC parser cannot match ⇒ the 4-bit override is silently dropped and loading dies later with

KeyError: 'layers.0.mlp.experts.routed_experts.w2_weight_packed'

Already applied in this directory (config.json.pristine keeps the stock export):

producer extra_config key INC matches it?
0.15.0 stock (this build, .pristine) .mlp.experts. ❌
this build (config.json, normalized) .*mlp\.experts ✅
0.16.0 stock (-MXFP4-Mixed-Tuning-AutoRound) .*mlp\.experts.* ✅ without any edit

3. On Hopper (SM90: H100/H200) this is a memory saving, not a speed saving

No FP4 tensor cores on SM90 ⇒ experts run through Marlin as weight-only FP4 (effectively W4A16):

Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE
Using MarlinMxfp8LinearKernel for MXFP8 GEMM
WARNING [inc_mxfp4_moe.py:204] This device lacks native FP4 compute; using weight-only FP4 via the Marlin kernel …

Dense Linear (here: shared_mlp) does get dynamic MXFP8 activation quantization; attention is now plain BF16 in this build.


Overview

Field Value
Base model tencent/HunyuanImage-3.0-Instruct-Distil (cfg_distilled=true, use_meanflow=true), local BF16 copy /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil
MoE geometry 32 layers × 64 routed experts, moe_topk=8, 1 shared expert/layer, hidden 4096, moe_intermediate=3072
Scheme Mixed: routed experts → MXFP4 (E2M1, group 32, E8M0); mlp.shared_mlp → MXFP8 (E4M3, group 32, E8M0); self_attn.* → BF16 (the delta vs -MXFP4-ar)
Export format --format auto_round → quant_method="auto-round", packing_format="auto_round:llm_compressor"
Loader path vLLM INC (INCConfig.override_quantization_method() maps "auto-round" → "inc")
Quantization tool auto-round 0.15.0, --model_free (no calibration), wall time 19 min 35 s on 1× GPU (08:04:22 → 08:23:57, exit=0)
Post-processing tools/normalize_experts_extra_config_key.py applied (see Read-this-first 2)
Disk size 53.89 GB (50.19 GiB) in 32 shards — -MXFP4-ar 52.59 GB, -MXFP8 91.25 GB, BF16 base 168.61 GB ⇒ 0.32× BF16
Parameters total 83.04 B; quantized 4 160 tensors = experts 77.31 B @4-bit + shared_mlp 1.21 B @8-bit = 94.55 % of parameters (-MXFP4-ar: 4 224 tensors / 96.17 %)
Tensor count 9 329 — vs 9 393 in -MXFP4-ar: the 64 attention weight_scale tensors are gone, everything else identical (torch.equal verified on experts/shared_mlp)
Kept in BF16 / FP32 self_attn.qkv_proj, self_attn.o_proj (1.34 B params — new in this build), all layernorms, the MoE router mlp.gate.wg (32), ViT + vision_aligner (BF16), VAE (FP32, as in base), wte/lm_head, guidance_emb/timestep_emb/timestep_r_emb/time_embed*/patch_embed/final_layer
Group size 32 everywhere, verified from tensor shapes (4 160 scale tensors), not read from config

Module-level coverage (measured from safetensors headers)

scripts/module_scheme_inventory.py classification: F8_E4M3 + same-name weight_scale ⇒ MXFP8 · U8 weight_packed ⇒ MXFP4 · no same-name scale ⇒ unquantized.

Backbone model.layers.0…31 (137 tensors/layer, 32 identical layers):

module template per layer this build -MXFP4-ar -MXFP8
self_attn.qkv_proj [6144,4096] 1 BF16 MXFP8 MXFP8
self_attn.o_proj [4096,4096] 1 BF16 MXFP8 MXFP8
self_attn.query_layernorm / key_layernorm [128] 2 BF16 BF16 BF16
mlp.experts.M.gate_and_up_proj → U8 [6144,2048] + scale U8 [6144,128] 64 MXFP4 MXFP4 MXFP8
mlp.experts.M.down_proj → U8 [4096,1536] + scale U8 [4096,96] 64 MXFP4 MXFP4 MXFP8
mlp.shared_mlp.gate_and_up_proj [6144,4096] 1 MXFP8 MXFP8 MXFP8
mlp.shared_mlp.down_proj [4096,3072] 1 MXFP8 MXFP8 MXFP8
mlp.gate.wg (MoE router) [64,4096] 1 BF16 BF16 BF16
input_layernorm / post_attention_layernorm [4096] 2 BF16 BF16 BF16

Outside the backbone — identical in all three builds, nothing quantized: ViT (vision_model.encoder.layers.0…26.{mlp.fc1 [4304,1152], mlp.fc2 [1152,4304], self_attn.{q,k,v,out}_proj, layer_norm1/2}, embeddings.*, head.*), vision_aligner.layers.{0,2}, VAE (vae.encoder.* / vae.decoder.*, FP32 convolutions), model.wte/lm_head, patch_embed/final_layer/time_embed*/ timestep_emb/timestep_r_emb/guidance_emb, model.ln_f.

Byte account, self-consistent to 0 MB: attention = 1.342 B params. MXFP8 costs 1.342 × (1 + 1/32) = 1.384 GB; BF16 costs 2.684 GB ⇒ predicted 52.59 + 1.300 = 53.89 GB, measured 53.89 GB.

quantization_config from config.json (verbatim, extra_config collapsed):

{
  "quant_method": "auto-round",
  "packing_format": "auto_round:llm_compressor",
  "bits": 8, "group_size": 32, "sym": true, "data_type": "mx_fp",
  "act_bits": 8, "act_data_type": "mx_fp", "act_dynamic": true, "act_group_size": 32, "act_sym": true,
  "model_free": true, "iters": 0, "enable_quanted_input": false,
  "autoround_version": "0.15.0",
  "block_name_to_quantize": "model.layers",
  "extra_config": {
    ".*mlp\\.experts":                      { "bits": 4 },                      // ← normalized (stock: ".mlp.experts.")
    "model.layers.N.self_attn.qkv_proj":    { "bits": 16, "data_type": "float", "act_bits": 16, "act_data_type": "float" },
    "model.layers.N.self_attn.o_proj":      { "bits": 16, … },                  // ← 64 new entries = this build's delta
    "... 220 more bits:16 entries (ViT / VAE-side / heads / router) ..."
  }
}

extra_config holds 285 entries: 1 at bits:4 + 284 at bits:16. Compared with -MXFP4-ar (221 entries) the diff is exactly +64 = 32 layers × {qkv_proj, o_proj}, with nothing removed. There is no ignore list in this format — 16-bit exceptions are extra_config entries — and everything outside model.layers.0…31 is untouched by construction (block_name_to_quantize).


Quantization command (expanded — no wrapper)

# ── 0) quantization environment: aqa venv's auto-round 0.15.0 (NOT the inference venv) ──
export OMP_NUM_THREADS=24
export AR_MODEL_FREE_SHARD_PARALLELISM=1          # caps model_free peak RAM (observed ~3-4 GB)
/home/kaokaolv/.venv-aqa-gpu/bin/python -c "import auto_round;print(auto_round.__version__)"   # -> 0.15.0

# ── 1) quantize ────────────────────────────────────────────────────────────
/home/kaokaolv/.venv-aqa-gpu/bin/auto-round \
  --model_name /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil \
  --model_free \
  --scheme MXFP8 \
  --layer_config '{.mlp.experts.:{bits:4,data_type:mx_fp}}' \
  --ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte,self_attn.qkv_proj,self_attn.o_proj" \
  --format auto_round \
  --device cuda:4 \
  --output_dir /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
  2>&1 | tee /home/kaokaolv/quantization_analysis/logs/quantize_mxfp4_ar_noattn.log

This is byte-for-byte what produced this checkpoint: start 2026-09-24 08:04:22, === exit=0 2026-09-24 08:23:57 ===, 19 min 35 s, one GPU. It differs from scripts/quantize_mxfp4_ar.sh (which produced -MXFP4-ar) only by the two extra self_attn.* entries in --ignore_layers.

Verify the ignore actually took effect, straight from the tool's own log:

grep -o "Ignored layers (3): .*" /home/kaokaolv/quantization_analysis/logs/quantize_mxfp4_ar_noattn.log | head -3
# Ignored layers (3): model.layers.0.mlp.gate.wg, model.layers.0.self_attn.o_proj, model.layers.0.self_attn.qkv_proj
# Ignored layers (3): model.layers.1.mlp.gate.wg, …        (once per layer, 32 layers)

2) Required follow-up: normalize the extra_config key

M=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn
mkdir -p $M/tools
cp -f /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar/tools/normalize_experts_extra_config_key.py $M/tools/
/home/kaokaolv/.venv-omini-latest/bin/python $M/tools/normalize_experts_extra_config_key.py $M
/home/kaokaolv/.venv-omini-latest/bin/python $M/tools/normalize_experts_extra_config_key.py $M --check
config.json: backup -> config.json.pristine
config.json: normalized ✅  '.mlp.experts.' -> '.*mlp\\.experts' = {'bits': 4}
quantization_config.json: normalized ✅  '.mlp.experts.' -> '.*mlp\\.experts' = {'bits': 4}

3) Acceptance checks (all four were run; expect exactly this output)

M=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn
PY=/home/kaokaolv/.venv-omini-latest/bin/python

# ① experts really are 4-bit packed
$PY -c "
import json,collections
wm=json.load(open('$M/model.safetensors.index.json'))['weight_map']
print(dict(collections.Counter(k.split('.')[-1] for k in wm if 'experts' in k)))"
#   {'weight_packed': 4096, 'weight_scale': 4096}

# ② attention really has NO scales  (use EXACT suffixes — see caveat below)
$PY -c "
import json
wm=json.load(open('$M/model.safetensors.index.json'))['weight_map']
for s in ('self_attn.qkv_proj','self_attn.o_proj','mlp.gate.wg'):
    print(f'{s}: weight={sum(1 for k in wm if k.endswith(s+\".weight\"))} '
          f'weight_scale={sum(1 for k in wm if k.endswith(s+\".weight_scale\"))} '
          f'weight_packed={sum(1 for k in wm if k.endswith(s+\".weight_packed\"))}')"
#   self_attn.qkv_proj: weight=32 weight_scale=0 weight_packed=0
#   self_attn.o_proj:   weight=32 weight_scale=0 weight_packed=0
#   mlp.gate.wg:        weight=32 weight_scale=0 weight_packed=0

# ③ coverage / group sizes, straight from headers (no model load)
$PY /home/kaokaolv/quantization_analysis/scripts/module_scheme_inventory.py $M --show-groups
#   张量(Linear 类) 4777 | 逻辑参数 83.040 B | 量化覆盖 78.517 B (94.55 %)
#   实测 group_size 分布: {32: 4160}

# ④ ask the REAL vLLM code how it dispatches (exit 0 = loadable)
cd /tmp && $PY /home/kaokaolv/quantization_analysis/scripts/check_autoround_experts_key.py $M
#   routed experts 容器   bits=4   INCMxfp4Scheme
#   shared_mlp            bits=8   INCMxfp8Scheme
#   self_attn.qkv_proj   bits=16   (未量化 → Unquantized*)
#   MoE router gate.wg   bits=16   (未量化 → Unquantized*)
#   ViT fc2              bits=16   (未量化 → Unquantized*)
#   ✅ experts 正确解析为 4bit,可直接推理

⚠️ Do not test ② with a substring. 'mlp.gate' in k also matches mlp.shared_**mlp.gate**_and_up_proj.weight_scale and yields the opposite (wrong) verdict. Always use k.endswith('<module>.weight_scale').

Command caveats worth repeating

  • Never put a bare gate in --ignore_layers — to_standard_regex expands it to .*gate.*, which also swallows the fused expert projections gate_and_up_proj (4096 tensors) ⇒ silently wrong coverage, no error. This build uses fully-qualified self_attn.qkv_proj / self_attn.o_proj.
  • The MoE router needs no entry: AutoRound hard-skips anything matching .gate. (auto_round/utils/model_free_utils.py:43, _BLOCK_NAME_TO_IGNORE).
  • vision must stay ignored: ViT mlp.fc2 has in_features=4304 (header: fc2.weight=[1152,4304]), not a multiple of the 32-element MX group ⇒ MXFP8 requires input_size_per_partition (2152) to be divisible by 32 at TP2.
  • The log line Scheme: … bits=8 / Packing format: mxfp8-quantized is the default scheme echo, not evidence that --layer_config was dropped. Judge 4-bit from headers (weight_packed count).
  • auto-round has no --version flag (unrecognized arguments); query the package instead.

Inference environment

Component Version
vLLM 0.29.0 — required: INC MXFP8 MoE method (INCMxfp8MoEMethod) exists only from 0.29
vLLM-Omni latest main @ 1c7476ec19899f1e61838e23e9fad3505403b9b9 (0.29.0rc2.dev161+g1c7476ec1), editable
PyTorch 2.13.0+cu132 · FlashInfer 0.6.18 · Python 3.12.3
GPU NVIDIA H200 141 GB — DiT-only 2×TP2, AR+DiT 2×(TP1+TP1), AR-only 1×

vLLM-Omni local patches are required for the auto_round/MXFP path (A/B/C/E/G/H/I/K/K2). Without them you get KeyError: '…w13_weight_scale', the 2152 divisible by 32 error, or — worst — silently degenerate images. Details: HunyuanImage-3.0-Instruct-Distil_vllm-omni本地补丁说明.md (in quantization_analysis/docs/).

# common environment for every mode below
V=/home/kaokaolv/.venv-omini-latest
OMNI=/home/kaokaolv/quantization_analysis/latest/vllm-omni
export PYTHONPATH=$OMNI
export CUDA_HOME=$V/lib/python3.12/site-packages/nvidia/cu13
export PATH=$CUDA_HOME/bin:$V/bin:$HOME/.local/bin:$PATH     # ninja lives in $V/bin; vLLM JIT needs it on PATH
export NCCL_NVLS_ENABLE=0              # this host's NVLS fabric is broken: TP>=2 group init hits NCCL error
export VLLM_USE_FLASHINFER_SAMPLER=0   # FlashInfer sampler JIT fails on this CUDA header set
unset  VLLM_ALLREDUCE_USE_FLASHINFER   # keep vLLM 0.29 default (fixed upstream by vllm-omni #7676)
mkdir -p /tmp/hy_run && cd /tmp/hy_run    # cwd must be neutral: children re-import vllm_omni from cwd

Deploy configs

Both files below are complete; save them and pass via --deploy-config. devices are local indices — they compose with CUDA_VISIBLE_DEVICES.

hunyuan_image3_dit_tp2.yaml (DiT-only, TP2):

pipeline: hunyuan_image_3_dit
async_chunk: false
trust_remote_code: true
stages:
  - stage_id: 0
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    trust_remote_code: true
    enforce_eager: true
    devices: "0,1"
    vae_use_slicing: false
    vae_use_tiling: false
    cache_backend:
    cache_config:
    enable_cache_dit_summary: false
    parallel_config:
      pipeline_parallel_size: 1
      data_parallel_size: 1
      tensor_parallel_size: 2
      enable_expert_parallel: true
      sequence_parallel_size: 1
      ulysses_degree: 1
      ring_degree: 1
      allgather_degree: 1
      cfg_parallel_size: 1
      vae_patch_parallel_size: 1
      use_hsdp: false
      hsdp_shard_size: -1
      hsdp_replicate_size: 1
    default_sampling_params:
      seed: 42

hunyuan_image_3_moe_2gpu_tp1.yaml (AR+DiT, 2 GPUs, AR TP1 + DiT TP1):

pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true
connectors:
  shared_memory_connector:
    name: SharedMemoryConnector
stages:
  - stage_id: 0                       # AR
    is_comprehension: true
    final_output: true
    final_output_type: text
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    max_num_batched_tokens: 32768
    devices: "0"
    tensor_parallel_size: 1
    hf_overrides:
      rope_parameters:
        mrope_section: [0, 32, 32]
        rope_type: default
    omni_kv_config:
      need_send_cache: true
    output_connectors:
      to_stage_1: shared_memory_connector
    default_sampling_params:
      temperature: 0.0                # greedy — see Fidelity note about AR determinism
      top_p: 1
      top_k: -1
      max_tokens: 8192
      detokenize: true
      skip_special_tokens: false
      include_stop_str_in_output: true
  - stage_id: 1                       # DiT
    max_num_seqs: 1
    gpu_memory_utilization: 0.9
    enforce_eager: true
    devices: "1"
    distributed_executor_backend: "mp"
    omni_kv_config:
      need_recv_cache: true
    parallel_config:
      tensor_parallel_size: 1
      enable_expert_parallel: true
    input_connectors:
      from_stage_0: shared_memory_connector
    default_sampling_params:
      num_inference_steps: 8
      guidance_scale: 0               # overridden by --guidance-scale below
edges:
  - from: 0
    to: 1
    window_size: -1
    max_inflight: 1

Mode 1 — DiT-only (2 GPUs, TP2). No AR stage at all.

export CUDA_VISIBLE_DEVICES=4,5
/home/kaokaolv/.venv-omini-latest/bin/python \
  /home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model                /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
  --deploy-config        ./hunyuan_image3_dit_tp2.yaml \
  --prompt               "A cute cat" \
  --num-inference-steps  8 \
  --guidance-scale       5.0 \
  --seed                 42 \
  --init-timeout         1800 \
  --stage-init-timeout   1500 \
  --output               ./dit_only_tp2.png

Observed (exit=0, 2026-09-24 08:42:51 → 08:47:31):

Stage 0 logical-to-physical device mapping: 0->0, 1->1
Using MarlinMxfp8LinearKernel for MXFP8 GEMM
Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE
Model loading took 25.3899 GiB and … seconds        # per card
image_pixels=1048576                                 # 1024x1024
Saved generated image to ./dit_only_tp2.png

The prompt goes straight to the DiT — grep -c ar2diffusion on this run's log is 0, and there is no target size= line (canvas = the example's default --height/--width = 1024).

Mode 2 — AR+DiT (2 GPUs, AR TP1 + DiT TP1)

export CUDA_VISIBLE_DEVICES=4,5
/home/kaokaolv/.venv-omini-latest/bin/python \
  /home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model                /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
  --deploy-config        ./hunyuan_image_3_moe_2gpu_tp1.yaml \
  --prompt               "A cute cat" \
  --num-inference-steps  8 \
  --guidance-scale       5.0 \
  --seed                 42 \
  --init-timeout         1800 \
  --stage-init-timeout   1500 \
  --output               ./ar_dit_full.png

Observed (exit=0, 08:47:31 → 08:54:07):

Stage 0 … 0->0   (AR)   Model loading took 47.67 GiB
Stage 1 … 1->1   (DiT)  Model loading took 47.8406 GiB
[ar2diffusion] Request 0: AR generated 12 tokens, text length=82, cot_text length=82,
               target size=576x1472 (AR ratio_idx=9)
Saved generated image to ./ar_dit_full.png

-MXFP4-ar loads at 46.69 / 46.63 GiB in the same slots ⇒ +≈1.0 GiB per stage, matching the BF16 attention. ⚠️ Note the AR branch (12 tokens, ratio_idx=9) — see Fidelity before comparing this image with any other build's full-pipeline output.

Mode 3 — AR-only (1 GPU, text out) — different example script

export CUDA_VISIBLE_DEVICES=4
/home/kaokaolv/.venv-omini-latest/bin/python \
  /home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/x_to_text/x_to_text.py \
  --model          /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
  --deploy-config  ./hunyuan_image3_ar_tp1.yaml \
  --trust-remote-code \
  --prompt         "A cute cat" \
  --max-tokens     256 \
  --output         ./ar_only.txt

⚠️ x_to_text.py does not accept --init-timeout / --stage-init-timeout (verified: argparse exit=2). Only text_to_image.py does.

Parameter notes

--num-inference-steps 8 (distilled checkpoint — do not use 50). --guidance-scale is a real input (cfg_distilled=true; vLLM-Omni feeds 1000 × guidance_scale as the guidance embedding) — the images here used 5.0; upstream's Distil e2e reference uses 2.5, so keep it fixed across any comparison. --prompt reaches only the AR stage in the two-stage config. Saved PNGs are always 1024×1024 — target size in the log is the AR's predicted conditioning aspect, not the canvas (text_to_image.py:794 saves without resizing; 17/17 logged runs verified).


Fidelity

Weight side — measured on this checkpoint (not borrowed from the sibling build)

Dequantizing this build and comparing with the local BF16 base (scripts/dequant_error_vs_base.py, relative L2):

Module scheme here rel. L2 vs BF16 -MXFP4-ar (attn@MXFP8)
self_attn.qkv_proj L0 / self_attn.o_proj L31 BF16 0.0000 0.0267
mlp.shared_mlp.gate_and_up_proj L0 / down_proj L31 MXFP8 0.0267 0.0267
mlp.experts.* (4 sampled: L0/L15/L31, both projections) MXFP4 0.1113 – 0.1125 0.1113 – 0.1125
mlp.gate.wg router BF16 0.0000 0.0000

Bit-level proof that the change is surgical (torch.equal on the raw tensors):

check result
this build's self_attn.qkv_proj / o_proj vs BF16 base torch.equal = True (dtype bfloat16, untouched copy)
this build's experts weight_packed / weight_scale, and shared_mlp.weight, vs -MXFP4-ar torch.equal = True

⇒ The only tensors that differ between this checkpoint and -MXFP4-ar are the 64 attention weights (and the 64 weight_scale tensors that no longer exist: 9 329 tensors here vs 9 393 there). That is also why the two builds' images agree at 32.97 dB / SSIM@1/8 0.9806: numerically they share every expert and shared-expert weight bit for bit.

Image side — DiT-only, all four builds in one session (2026-09-24 08:27–08:48, cards 4/5, TP2)

DiT-only four-way

cell build PSNR vs BF16 SSIM@1 SSIM@1/8 std
① BF16 base (reference) — — — 47.37
② -MXFP8 (all 8-bit, INC) 32.77 0.9760 0.9813 48.30
③ -MXFP4-ar (attn@MXFP8) 19.51 0.8421 0.6627 48.77
④ this build (attn@BF16) 19.60 0.8494 0.6699 48.22

③ ↔ ④ = 32.97 dB / SSIM@1/8 0.9806 — attention precision is not the driver. The MXFP4 gap vs BF16 (0.98 → 0.66 as the SSIM scale gets coarser ⇒ composition-level, not texture) is attributable to the experts' 4-bit codebook.

Session control: this session's BF16 DiT-only capture is md5-identical to the 2026-09-22 capture of the same config (fc26e6ee66277ff8b91dfa0d3cfaaf7e), so BF16 DiT-only is bit-reproducible across days here. Cross-day MXFP8-ar vs itself is 32.77 dB (mild numeric jitter, same composition).

AR+DiT — ⚠️ the three runs took three different AR branches

AR+DiT branches

build AR output (verbatim) tokens conditioning aspect
-MXFP8 It's a<boi><img_size_1024><img_ratio_14> 6 896×1152
-MXFP4-ar It's a warm and healing style<boi><img_size_1024><img_ratio_14> 10 896×1152
this build sitting upright on a light-colored wooden floor.<boi><img_size_1024><img_ratio_9> 12 576×1472

The AR output is the DiT conditioning, so these images are not pixel-comparable: vs -MXFP8 this build scores 15.14 dB / SSIM@1/8 0.4383 and -MXFP4-ar scores 18.52 / 0.4976 — those numbers measure branch divergence, not quantization fidelity. Rule: before any full-pipeline comparison, check that [ar2diffusion] AR generated N tokens and the AR text match on both sides. Use the DiT-only figure for fidelity judgments.

Individual outputs, uncropped:

BF16 (DiT TP2) -MXFP8 (DiT TP2) -MXFP4-ar (DiT TP2) this build (DiT TP2)
a b c d

AR-only text is not cross-comparable

images/ar_only_output.txt (1168 bytes, full description, exit=0) is one of only two long captures in this workspace; the same command gave 7-byte stubs (It's a) for BF16, -MXFP8 and -MXFP4-ar, and 27-30 bytes for other builds, and one build even produced both 7 B and 1198 B on repeat runs. The x_to_text.py path is itself unstable; treat AR-only text as "the AR path loads and generates", never as evidence of a model difference.

Not measured

No task-level benchmark (GenEval / DPG-Bench / CVTG-2K / DrawBench / WISE) has been run for this build — nor for -MXFP4-ar / -MXFP4-mixed / the tuned build. Only BF16 and -MXFP8-ct have full GenEval numbers in this workspace. "Composition diverges from BF16" is measured; "quality is worse" is not established, and all image evidence here rests on a single prompt (A cute cat).


Known limitations

  1. Requires the metadata fix (Read-this-first 2); fails loudly without it — which is preferable to loading with the wrong scheme.
  2. The self_attn → BF16 change is not a fidelity win (32.97 dB against the attn-quantized sibling, both ~0.67 SSIM@1/8 vs BF16): +1.30 GB disk, +1.0 GiB/GPU, no measurable improvement.
  3. Memory-only on Hopper — experts run W4A16 through Marlin.
  4. Composition diverges from BF16/MXFP8 (SSIM@1/8 ≈0.67, DiT-only same-session) — caused by MXFP4 experts.
  5. ViT stays BF16 (vision_model.encoder.*.mlp.fc2 input dim 4304 % 32 ≠ 0) ⇒ not a whole-model MXFP4 build.
  6. Full-pipeline runs are not comparable across builds unless the AR branch matches; this build landed on ratio_idx=9 where BF16/MXFP8 use ratio_idx=14.
  7. block_name_to_quantize = model.layers scopes quantization to the 32 shared backbone layers; AR and DiT load the same tensors through two separate stacks (weights are not shared between stages).
  8. images/ holds verification snapshots only; nothing there is read when loading weights.

Contents

README.md                                       this file
config.json                                     [NORMALIZED] quant_method=auto-round, extra_config (285 entries)
config.json.pristine                            stock 0.15.0 export (key ".mlp.experts.") — fails to load
quantization_config.json                        [NORMALIZED] secondary discovery carrier
model-000NN-of-00032.safetensors        53.89 GB total, **9329** tensors
                                        (= 9393 in -MXFP4-ar minus the 64 attention weight_scale tensors)
model.safetensors.index.json
tools/normalize_experts_extra_config_key.py     apply / --check / --revert the metadata fix
images/
  ar_dit_full_2gpu.png                          AR+DiT, 2 GPUs (hero; AR branch ratio_idx=9)
  dit_only_tp2.png                              this build, DiT-only TP2
  bf16_dit_only_reference.png                   BF16 base, DiT-only TP2 (same session)
  mxfp8_ar_dit_only.png                         -MXFP8 (auto_round), DiT-only TP2 (same session)
  sibling_mxfp4_ar_dit_only.png                 -MXFP4-ar, DiT-only TP2 (same session)
  grid_ditonly_fourway.png                      BF16 / MXFP8 / MXFP4-ar / this build + PSNR+SSIM labels
  grid_ar_dit_branches.png                      three AR+DiT outputs, each labelled with its AR branch
  ar_only_output.txt                            AR-only text (1168 B; see the AR-only caveat)
*.py / tokenizer* / utils/                      inherited from the base model

Reproduce the analysis tables with quantization_analysis/scripts/module_scheme_inventory.py <dir> --show-groups (archived outputs: quantization_analysis/reference/scheme_templates/).

Downloads last month
10
Safetensors
Model size
44B params
Tensor type
BF16
·
U8
·
F8_E4M3
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support