YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
HunyuanImage-3.0-Instruct-Distil · MXFP4 mixed-precision with attention kept in BF16 (auto-round / vLLM INC path)
Same MXFP4-for-routed-experts recipe as
-MXFP4-ar, plus one deliberate change:
self_attn.qkv_proj and self_attn.o_proj are added to --ignore_layers, so attention stays BF16.
Recipe summary: routed experts → MXFP4 · mlp.shared_mlp → MXFP8 · self_attn → BF16 ·
everything else (ViT, VAE, MoE router, embeddings, DiT heads) → BF16/FP32.
Generated output, not a ground-truth reference. AR+DiT on 2 GPUs (AR TP1 + DiT TP1), prompt
A cute cat,seed=42, 8 steps,guidance_scale=5.0. Captured 2026-09-24 08:54,exit=0.
⚠️ Read this first
1. This build is -MXFP4-ar + "un-quantize attention" — and it buys you almost nothing
Measured in one session, same topology, DiT-only TP2 (so attention precision is the only variable):
| pair | PSNR | SSIM@1 | SSIM@1/8 |
|---|---|---|---|
this build ↔ -MXFP4-ar |
32.97 | 0.9802 | 0.9806 |
| this build ↔ BF16 | 19.60 | 0.8494 | 0.6699 |
-MXFP4-ar ↔ BF16 |
19.51 | 0.8421 | 0.6627 |
-MXFP8 (all 8-bit) ↔ BF16 |
32.77 | 0.9760 | 0.9813 |
⇒ Restoring BF16 attention costs +1.30 GB on disk and +≈1.0 GiB/GPU at load time, and changes the picture at the level of run-to-run numerical noise. The composition divergence of MXFP4 comes from the experts' 4-bit codebook, not from attention precision. Keep this build only if you have a specific reason to protect attention (e.g. a downstream experiment), not as a fidelity fix.
2. A stock --format auto_round export of this recipe does NOT load — the metadata fix is required
AutoRound 0.15.0 writes the --layer_config pattern into extra_config verbatim
(key = ".mlp.experts."), which vLLM's INC parser cannot match ⇒ the 4-bit override is silently dropped
and loading dies later with
KeyError: 'layers.0.mlp.experts.routed_experts.w2_weight_packed'
Already applied in this directory (config.json.pristine keeps the stock export):
| producer | extra_config key |
INC matches it? |
|---|---|---|
0.15.0 stock (this build, .pristine) |
.mlp.experts. |
❌ |
this build (config.json, normalized) |
.*mlp\.experts |
✅ |
0.16.0 stock (-MXFP4-Mixed-Tuning-AutoRound) |
.*mlp\.experts.* |
✅ without any edit |
3. On Hopper (SM90: H100/H200) this is a memory saving, not a speed saving
No FP4 tensor cores on SM90 ⇒ experts run through Marlin as weight-only FP4 (effectively W4A16):
Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE
Using MarlinMxfp8LinearKernel for MXFP8 GEMM
WARNING [inc_mxfp4_moe.py:204] This device lacks native FP4 compute; using weight-only FP4 via the Marlin kernel …
Dense Linear (here: shared_mlp) does get dynamic MXFP8 activation quantization; attention is now plain
BF16 in this build.
Overview
| Field | Value |
|---|---|
| Base model | tencent/HunyuanImage-3.0-Instruct-Distil (cfg_distilled=true, use_meanflow=true), local BF16 copy /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil |
| MoE geometry | 32 layers × 64 routed experts, moe_topk=8, 1 shared expert/layer, hidden 4096, moe_intermediate=3072 |
| Scheme | Mixed: routed experts → MXFP4 (E2M1, group 32, E8M0); mlp.shared_mlp → MXFP8 (E4M3, group 32, E8M0); self_attn.* → BF16 (the delta vs -MXFP4-ar) |
| Export format | --format auto_round → quant_method="auto-round", packing_format="auto_round:llm_compressor" |
| Loader path | vLLM INC (INCConfig.override_quantization_method() maps "auto-round" → "inc") |
| Quantization tool | auto-round 0.15.0, --model_free (no calibration), wall time 19 min 35 s on 1× GPU (08:04:22 → 08:23:57, exit=0) |
| Post-processing | tools/normalize_experts_extra_config_key.py applied (see Read-this-first 2) |
| Disk size | 53.89 GB (50.19 GiB) in 32 shards — -MXFP4-ar 52.59 GB, -MXFP8 91.25 GB, BF16 base 168.61 GB ⇒ 0.32× BF16 |
| Parameters | total 83.04 B; quantized 4 160 tensors = experts 77.31 B @4-bit + shared_mlp 1.21 B @8-bit = 94.55 % of parameters (-MXFP4-ar: 4 224 tensors / 96.17 %) |
| Tensor count | 9 329 — vs 9 393 in -MXFP4-ar: the 64 attention weight_scale tensors are gone, everything else identical (torch.equal verified on experts/shared_mlp) |
| Kept in BF16 / FP32 | self_attn.qkv_proj, self_attn.o_proj (1.34 B params — new in this build), all layernorms, the MoE router mlp.gate.wg (32), ViT + vision_aligner (BF16), VAE (FP32, as in base), wte/lm_head, guidance_emb/timestep_emb/timestep_r_emb/time_embed*/patch_embed/final_layer |
| Group size | 32 everywhere, verified from tensor shapes (4 160 scale tensors), not read from config |
Module-level coverage (measured from safetensors headers)
scripts/module_scheme_inventory.py classification: F8_E4M3 + same-name weight_scale ⇒ MXFP8 ·
U8 weight_packed ⇒ MXFP4 · no same-name scale ⇒ unquantized.
Backbone model.layers.0…31 (137 tensors/layer, 32 identical layers):
| module template | per layer | this build | -MXFP4-ar |
-MXFP8 |
|---|---|---|---|---|
self_attn.qkv_proj [6144,4096] |
1 | BF16 | MXFP8 | MXFP8 |
self_attn.o_proj [4096,4096] |
1 | BF16 | MXFP8 | MXFP8 |
self_attn.query_layernorm / key_layernorm [128] |
2 | BF16 | BF16 | BF16 |
mlp.experts.M.gate_and_up_proj → U8 [6144,2048] + scale U8 [6144,128] |
64 | MXFP4 | MXFP4 | MXFP8 |
mlp.experts.M.down_proj → U8 [4096,1536] + scale U8 [4096,96] |
64 | MXFP4 | MXFP4 | MXFP8 |
mlp.shared_mlp.gate_and_up_proj [6144,4096] |
1 | MXFP8 | MXFP8 | MXFP8 |
mlp.shared_mlp.down_proj [4096,3072] |
1 | MXFP8 | MXFP8 | MXFP8 |
mlp.gate.wg (MoE router) [64,4096] |
1 | BF16 | BF16 | BF16 |
input_layernorm / post_attention_layernorm [4096] |
2 | BF16 | BF16 | BF16 |
Outside the backbone — identical in all three builds, nothing quantized: ViT
(vision_model.encoder.layers.0…26.{mlp.fc1 [4304,1152], mlp.fc2 [1152,4304], self_attn.{q,k,v,out}_proj, layer_norm1/2}, embeddings.*, head.*), vision_aligner.layers.{0,2}, VAE (vae.encoder.* /
vae.decoder.*, FP32 convolutions), model.wte/lm_head, patch_embed/final_layer/time_embed*/
timestep_emb/timestep_r_emb/guidance_emb, model.ln_f.
Byte account, self-consistent to 0 MB: attention = 1.342 B params. MXFP8 costs
1.342 × (1 + 1/32) = 1.384 GB; BF16 costs 2.684 GB ⇒ predicted 52.59 + 1.300 = 53.89 GB,
measured 53.89 GB.
quantization_config from config.json (verbatim, extra_config collapsed):
{
"quant_method": "auto-round",
"packing_format": "auto_round:llm_compressor",
"bits": 8, "group_size": 32, "sym": true, "data_type": "mx_fp",
"act_bits": 8, "act_data_type": "mx_fp", "act_dynamic": true, "act_group_size": 32, "act_sym": true,
"model_free": true, "iters": 0, "enable_quanted_input": false,
"autoround_version": "0.15.0",
"block_name_to_quantize": "model.layers",
"extra_config": {
".*mlp\\.experts": { "bits": 4 }, // ← normalized (stock: ".mlp.experts.")
"model.layers.N.self_attn.qkv_proj": { "bits": 16, "data_type": "float", "act_bits": 16, "act_data_type": "float" },
"model.layers.N.self_attn.o_proj": { "bits": 16, … }, // ← 64 new entries = this build's delta
"... 220 more bits:16 entries (ViT / VAE-side / heads / router) ..."
}
}
extra_config holds 285 entries: 1 at bits:4 + 284 at bits:16. Compared with -MXFP4-ar
(221 entries) the diff is exactly +64 = 32 layers × {qkv_proj, o_proj}, with nothing removed.
There is no ignore list in this format — 16-bit exceptions are extra_config entries — and everything
outside model.layers.0…31 is untouched by construction (block_name_to_quantize).
Quantization command (expanded — no wrapper)
# ── 0) quantization environment: aqa venv's auto-round 0.15.0 (NOT the inference venv) ──
export OMP_NUM_THREADS=24
export AR_MODEL_FREE_SHARD_PARALLELISM=1 # caps model_free peak RAM (observed ~3-4 GB)
/home/kaokaolv/.venv-aqa-gpu/bin/python -c "import auto_round;print(auto_round.__version__)" # -> 0.15.0
# ── 1) quantize ────────────────────────────────────────────────────────────
/home/kaokaolv/.venv-aqa-gpu/bin/auto-round \
--model_name /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil \
--model_free \
--scheme MXFP8 \
--layer_config '{.mlp.experts.:{bits:4,data_type:mx_fp}}' \
--ignore_layers "vision,guidance_emb,timestep_emb,timestep_r_emb,final_layer,wte,self_attn.qkv_proj,self_attn.o_proj" \
--format auto_round \
--device cuda:4 \
--output_dir /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
2>&1 | tee /home/kaokaolv/quantization_analysis/logs/quantize_mxfp4_ar_noattn.log
This is byte-for-byte what produced this checkpoint: start 2026-09-24 08:04:22, === exit=0 2026-09-24 08:23:57 ===, 19 min 35 s, one GPU. It differs from
scripts/quantize_mxfp4_ar.sh (which produced -MXFP4-ar) only by the two extra
self_attn.* entries in --ignore_layers.
Verify the ignore actually took effect, straight from the tool's own log:
grep -o "Ignored layers (3): .*" /home/kaokaolv/quantization_analysis/logs/quantize_mxfp4_ar_noattn.log | head -3
# Ignored layers (3): model.layers.0.mlp.gate.wg, model.layers.0.self_attn.o_proj, model.layers.0.self_attn.qkv_proj
# Ignored layers (3): model.layers.1.mlp.gate.wg, … (once per layer, 32 layers)
2) Required follow-up: normalize the extra_config key
M=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn
mkdir -p $M/tools
cp -f /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar/tools/normalize_experts_extra_config_key.py $M/tools/
/home/kaokaolv/.venv-omini-latest/bin/python $M/tools/normalize_experts_extra_config_key.py $M
/home/kaokaolv/.venv-omini-latest/bin/python $M/tools/normalize_experts_extra_config_key.py $M --check
config.json: backup -> config.json.pristine
config.json: normalized ✅ '.mlp.experts.' -> '.*mlp\\.experts' = {'bits': 4}
quantization_config.json: normalized ✅ '.mlp.experts.' -> '.*mlp\\.experts' = {'bits': 4}
3) Acceptance checks (all four were run; expect exactly this output)
M=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn
PY=/home/kaokaolv/.venv-omini-latest/bin/python
# ① experts really are 4-bit packed
$PY -c "
import json,collections
wm=json.load(open('$M/model.safetensors.index.json'))['weight_map']
print(dict(collections.Counter(k.split('.')[-1] for k in wm if 'experts' in k)))"
# {'weight_packed': 4096, 'weight_scale': 4096}
# ② attention really has NO scales (use EXACT suffixes — see caveat below)
$PY -c "
import json
wm=json.load(open('$M/model.safetensors.index.json'))['weight_map']
for s in ('self_attn.qkv_proj','self_attn.o_proj','mlp.gate.wg'):
print(f'{s}: weight={sum(1 for k in wm if k.endswith(s+\".weight\"))} '
f'weight_scale={sum(1 for k in wm if k.endswith(s+\".weight_scale\"))} '
f'weight_packed={sum(1 for k in wm if k.endswith(s+\".weight_packed\"))}')"
# self_attn.qkv_proj: weight=32 weight_scale=0 weight_packed=0
# self_attn.o_proj: weight=32 weight_scale=0 weight_packed=0
# mlp.gate.wg: weight=32 weight_scale=0 weight_packed=0
# ③ coverage / group sizes, straight from headers (no model load)
$PY /home/kaokaolv/quantization_analysis/scripts/module_scheme_inventory.py $M --show-groups
# 张量(Linear 类) 4777 | 逻辑参数 83.040 B | 量化覆盖 78.517 B (94.55 %)
# 实测 group_size 分布: {32: 4160}
# ④ ask the REAL vLLM code how it dispatches (exit 0 = loadable)
cd /tmp && $PY /home/kaokaolv/quantization_analysis/scripts/check_autoround_experts_key.py $M
# routed experts 容器 bits=4 INCMxfp4Scheme
# shared_mlp bits=8 INCMxfp8Scheme
# self_attn.qkv_proj bits=16 (未量化 → Unquantized*)
# MoE router gate.wg bits=16 (未量化 → Unquantized*)
# ViT fc2 bits=16 (未量化 → Unquantized*)
# ✅ experts 正确解析为 4bit,可直接推理
⚠️ Do not test ② with a substring. 'mlp.gate' in k also matches
mlp.shared_**mlp.gate**_and_up_proj.weight_scale and yields the opposite (wrong) verdict.
Always use k.endswith('<module>.weight_scale').
Command caveats worth repeating
- Never put a bare
gatein--ignore_layers—to_standard_regexexpands it to.*gate.*, which also swallows the fused expert projectionsgate_and_up_proj(4096 tensors) ⇒ silently wrong coverage, no error. This build uses fully-qualifiedself_attn.qkv_proj/self_attn.o_proj. - The MoE router needs no entry: AutoRound hard-skips anything matching
.gate.(auto_round/utils/model_free_utils.py:43,_BLOCK_NAME_TO_IGNORE). visionmust stay ignored: ViTmlp.fc2hasin_features=4304(header:fc2.weight=[1152,4304]), not a multiple of the 32-element MX group ⇒MXFP8 requires input_size_per_partition (2152) to be divisible by 32at TP2.- The log line
Scheme: … bits=8/Packing format: mxfp8-quantizedis the default scheme echo, not evidence that--layer_configwas dropped. Judge 4-bit from headers (weight_packedcount). auto-roundhas no--versionflag (unrecognized arguments); query the package instead.
Inference environment
| Component | Version |
|---|---|
| vLLM | 0.29.0 — required: INC MXFP8 MoE method (INCMxfp8MoEMethod) exists only from 0.29 |
| vLLM-Omni | latest main @ 1c7476ec19899f1e61838e23e9fad3505403b9b9 (0.29.0rc2.dev161+g1c7476ec1), editable |
| PyTorch | 2.13.0+cu132 · FlashInfer 0.6.18 · Python 3.12.3 |
| GPU | NVIDIA H200 141 GB — DiT-only 2×TP2, AR+DiT 2×(TP1+TP1), AR-only 1× |
vLLM-Omni local patches are required for the auto_round/MXFP path (A/B/C/E/G/H/I/K/K2). Without
them you get KeyError: '…w13_weight_scale', the 2152 divisible by 32 error, or — worst — silently
degenerate images. Details: HunyuanImage-3.0-Instruct-Distil_vllm-omni本地补丁说明.md
(in quantization_analysis/docs/).
# common environment for every mode below
V=/home/kaokaolv/.venv-omini-latest
OMNI=/home/kaokaolv/quantization_analysis/latest/vllm-omni
export PYTHONPATH=$OMNI
export CUDA_HOME=$V/lib/python3.12/site-packages/nvidia/cu13
export PATH=$CUDA_HOME/bin:$V/bin:$HOME/.local/bin:$PATH # ninja lives in $V/bin; vLLM JIT needs it on PATH
export NCCL_NVLS_ENABLE=0 # this host's NVLS fabric is broken: TP>=2 group init hits NCCL error
export VLLM_USE_FLASHINFER_SAMPLER=0 # FlashInfer sampler JIT fails on this CUDA header set
unset VLLM_ALLREDUCE_USE_FLASHINFER # keep vLLM 0.29 default (fixed upstream by vllm-omni #7676)
mkdir -p /tmp/hy_run && cd /tmp/hy_run # cwd must be neutral: children re-import vllm_omni from cwd
Deploy configs
Both files below are complete; save them and pass via --deploy-config.
devices are local indices — they compose with CUDA_VISIBLE_DEVICES.
hunyuan_image3_dit_tp2.yaml (DiT-only, TP2):
pipeline: hunyuan_image_3_dit
async_chunk: false
trust_remote_code: true
stages:
- stage_id: 0
max_num_seqs: 1
gpu_memory_utilization: 0.9
trust_remote_code: true
enforce_eager: true
devices: "0,1"
vae_use_slicing: false
vae_use_tiling: false
cache_backend:
cache_config:
enable_cache_dit_summary: false
parallel_config:
pipeline_parallel_size: 1
data_parallel_size: 1
tensor_parallel_size: 2
enable_expert_parallel: true
sequence_parallel_size: 1
ulysses_degree: 1
ring_degree: 1
allgather_degree: 1
cfg_parallel_size: 1
vae_patch_parallel_size: 1
use_hsdp: false
hsdp_shard_size: -1
hsdp_replicate_size: 1
default_sampling_params:
seed: 42
hunyuan_image_3_moe_2gpu_tp1.yaml (AR+DiT, 2 GPUs, AR TP1 + DiT TP1):
pipeline: hunyuan_image_3_moe
async_chunk: false
trust_remote_code: true
connectors:
shared_memory_connector:
name: SharedMemoryConnector
stages:
- stage_id: 0 # AR
is_comprehension: true
final_output: true
final_output_type: text
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
max_num_batched_tokens: 32768
devices: "0"
tensor_parallel_size: 1
hf_overrides:
rope_parameters:
mrope_section: [0, 32, 32]
rope_type: default
omni_kv_config:
need_send_cache: true
output_connectors:
to_stage_1: shared_memory_connector
default_sampling_params:
temperature: 0.0 # greedy — see Fidelity note about AR determinism
top_p: 1
top_k: -1
max_tokens: 8192
detokenize: true
skip_special_tokens: false
include_stop_str_in_output: true
- stage_id: 1 # DiT
max_num_seqs: 1
gpu_memory_utilization: 0.9
enforce_eager: true
devices: "1"
distributed_executor_backend: "mp"
omni_kv_config:
need_recv_cache: true
parallel_config:
tensor_parallel_size: 1
enable_expert_parallel: true
input_connectors:
from_stage_0: shared_memory_connector
default_sampling_params:
num_inference_steps: 8
guidance_scale: 0 # overridden by --guidance-scale below
edges:
- from: 0
to: 1
window_size: -1
max_inflight: 1
Mode 1 — DiT-only (2 GPUs, TP2). No AR stage at all.
export CUDA_VISIBLE_DEVICES=4,5
/home/kaokaolv/.venv-omini-latest/bin/python \
/home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
--deploy-config ./hunyuan_image3_dit_tp2.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 \
--guidance-scale 5.0 \
--seed 42 \
--init-timeout 1800 \
--stage-init-timeout 1500 \
--output ./dit_only_tp2.png
Observed (exit=0, 2026-09-24 08:42:51 → 08:47:31):
Stage 0 logical-to-physical device mapping: 0->0, 1->1
Using MarlinMxfp8LinearKernel for MXFP8 GEMM
Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE
Model loading took 25.3899 GiB and … seconds # per card
image_pixels=1048576 # 1024x1024
Saved generated image to ./dit_only_tp2.png
The prompt goes straight to the DiT — grep -c ar2diffusion on this run's log is 0, and there is no
target size= line (canvas = the example's default --height/--width = 1024).
Mode 2 — AR+DiT (2 GPUs, AR TP1 + DiT TP1)
export CUDA_VISIBLE_DEVICES=4,5
/home/kaokaolv/.venv-omini-latest/bin/python \
/home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
--deploy-config ./hunyuan_image_3_moe_2gpu_tp1.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 \
--guidance-scale 5.0 \
--seed 42 \
--init-timeout 1800 \
--stage-init-timeout 1500 \
--output ./ar_dit_full.png
Observed (exit=0, 08:47:31 → 08:54:07):
Stage 0 … 0->0 (AR) Model loading took 47.67 GiB
Stage 1 … 1->1 (DiT) Model loading took 47.8406 GiB
[ar2diffusion] Request 0: AR generated 12 tokens, text length=82, cot_text length=82,
target size=576x1472 (AR ratio_idx=9)
Saved generated image to ./ar_dit_full.png
-MXFP4-ar loads at 46.69 / 46.63 GiB in the same slots ⇒ +≈1.0 GiB per stage, matching the BF16
attention. ⚠️ Note the AR branch (12 tokens, ratio_idx=9) — see Fidelity before comparing this image
with any other build's full-pipeline output.
Mode 3 — AR-only (1 GPU, text out) — different example script
export CUDA_VISIBLE_DEVICES=4
/home/kaokaolv/.venv-omini-latest/bin/python \
/home/kaokaolv/quantization_analysis/latest/vllm-omni/examples/offline_inference/x_to_text/x_to_text.py \
--model /home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar-noattn \
--deploy-config ./hunyuan_image3_ar_tp1.yaml \
--trust-remote-code \
--prompt "A cute cat" \
--max-tokens 256 \
--output ./ar_only.txt
⚠️ x_to_text.py does not accept --init-timeout / --stage-init-timeout (verified: argparse
exit=2). Only text_to_image.py does.
Parameter notes
--num-inference-steps 8 (distilled checkpoint — do not use 50). --guidance-scale is a real input
(cfg_distilled=true; vLLM-Omni feeds 1000 × guidance_scale as the guidance embedding) — the images here
used 5.0; upstream's Distil e2e reference uses 2.5, so keep it fixed across any comparison.
--prompt reaches only the AR stage in the two-stage config. Saved PNGs are always 1024×1024 —
target size in the log is the AR's predicted conditioning aspect, not the canvas
(text_to_image.py:794 saves without resizing; 17/17 logged runs verified).
Fidelity
Weight side — measured on this checkpoint (not borrowed from the sibling build)
Dequantizing this build and comparing with the local BF16 base
(scripts/dequant_error_vs_base.py, relative L2):
| Module | scheme here | rel. L2 vs BF16 | -MXFP4-ar (attn@MXFP8) |
|---|---|---|---|
self_attn.qkv_proj L0 / self_attn.o_proj L31 |
BF16 | 0.0000 | 0.0267 |
mlp.shared_mlp.gate_and_up_proj L0 / down_proj L31 |
MXFP8 | 0.0267 | 0.0267 |
mlp.experts.* (4 sampled: L0/L15/L31, both projections) |
MXFP4 | 0.1113 – 0.1125 | 0.1113 – 0.1125 |
mlp.gate.wg router |
BF16 | 0.0000 | 0.0000 |
Bit-level proof that the change is surgical (torch.equal on the raw tensors):
| check | result |
|---|---|
this build's self_attn.qkv_proj / o_proj vs BF16 base |
torch.equal = True (dtype bfloat16, untouched copy) |
this build's experts weight_packed / weight_scale, and shared_mlp.weight, vs -MXFP4-ar |
torch.equal = True |
⇒ The only tensors that differ between this checkpoint and -MXFP4-ar are the 64 attention weights
(and the 64 weight_scale tensors that no longer exist: 9 329 tensors here vs 9 393 there).
That is also why the two builds' images agree at 32.97 dB / SSIM@1/8 0.9806: numerically they share every
expert and shared-expert weight bit for bit.
Image side — DiT-only, all four builds in one session (2026-09-24 08:27–08:48, cards 4/5, TP2)
| cell | build | PSNR vs BF16 | SSIM@1 | SSIM@1/8 | std |
|---|---|---|---|---|---|
| ① | BF16 base (reference) | — | — | — | 47.37 |
| ② | -MXFP8 (all 8-bit, INC) |
32.77 | 0.9760 | 0.9813 | 48.30 |
| ③ | -MXFP4-ar (attn@MXFP8) |
19.51 | 0.8421 | 0.6627 | 48.77 |
| ④ | this build (attn@BF16) | 19.60 | 0.8494 | 0.6699 | 48.22 |
③ ↔ ④ = 32.97 dB / SSIM@1/8 0.9806 — attention precision is not the driver. The MXFP4 gap vs BF16 (0.98 → 0.66 as the SSIM scale gets coarser ⇒ composition-level, not texture) is attributable to the experts' 4-bit codebook.
Session control: this session's BF16 DiT-only capture is md5-identical to the 2026-09-22 capture of the
same config (fc26e6ee66277ff8b91dfa0d3cfaaf7e), so BF16 DiT-only is bit-reproducible across days here.
Cross-day MXFP8-ar vs itself is 32.77 dB (mild numeric jitter, same composition).
AR+DiT — ⚠️ the three runs took three different AR branches
| build | AR output (verbatim) | tokens | conditioning aspect |
|---|---|---|---|
-MXFP8 |
It's a<boi><img_size_1024><img_ratio_14> |
6 | 896×1152 |
-MXFP4-ar |
It's a warm and healing style<boi><img_size_1024><img_ratio_14> |
10 | 896×1152 |
| this build | sitting upright on a light-colored wooden floor.<boi><img_size_1024><img_ratio_9> |
12 | 576×1472 |
The AR output is the DiT conditioning, so these images are not pixel-comparable: vs -MXFP8 this build
scores 15.14 dB / SSIM@1/8 0.4383 and -MXFP4-ar scores 18.52 / 0.4976 — those numbers measure branch
divergence, not quantization fidelity. Rule: before any full-pipeline comparison, check that
[ar2diffusion] AR generated N tokens and the AR text match on both sides. Use the DiT-only figure for
fidelity judgments.
Individual outputs, uncropped:
AR-only text is not cross-comparable
images/ar_only_output.txt (1168 bytes, full description, exit=0) is one of only two long captures in
this workspace; the same command gave 7-byte stubs (It's a) for BF16, -MXFP8 and -MXFP4-ar, and 27-30
bytes for other builds, and one build even produced both 7 B and 1198 B on repeat runs. The x_to_text.py
path is itself unstable; treat AR-only text as "the AR path loads and generates", never as evidence of a
model difference.
Not measured
No task-level benchmark (GenEval / DPG-Bench / CVTG-2K / DrawBench / WISE) has been run for this
build — nor for -MXFP4-ar / -MXFP4-mixed / the tuned build. Only BF16 and -MXFP8-ct have full GenEval
numbers in this workspace. "Composition diverges from BF16" is measured; "quality is worse" is not
established, and all image evidence here rests on a single prompt (A cute cat).
Known limitations
- Requires the metadata fix (Read-this-first 2); fails loudly without it — which is preferable to loading with the wrong scheme.
- The
self_attn→ BF16 change is not a fidelity win (32.97 dB against the attn-quantized sibling, both ~0.67 SSIM@1/8 vs BF16): +1.30 GB disk, +1.0 GiB/GPU, no measurable improvement. - Memory-only on Hopper — experts run W4A16 through Marlin.
- Composition diverges from BF16/MXFP8 (SSIM@1/8 ≈0.67, DiT-only same-session) — caused by MXFP4 experts.
- ViT stays BF16 (
vision_model.encoder.*.mlp.fc2input dim 4304 % 32 ≠ 0) ⇒ not a whole-model MXFP4 build. - Full-pipeline runs are not comparable across builds unless the AR branch matches; this build landed
on
ratio_idx=9where BF16/MXFP8 useratio_idx=14. block_name_to_quantize = model.layersscopes quantization to the 32 shared backbone layers; AR and DiT load the same tensors through two separate stacks (weights are not shared between stages).images/holds verification snapshots only; nothing there is read when loading weights.
Contents
README.md this file
config.json [NORMALIZED] quant_method=auto-round, extra_config (285 entries)
config.json.pristine stock 0.15.0 export (key ".mlp.experts.") — fails to load
quantization_config.json [NORMALIZED] secondary discovery carrier
model-000NN-of-00032.safetensors 53.89 GB total, **9329** tensors
(= 9393 in -MXFP4-ar minus the 64 attention weight_scale tensors)
model.safetensors.index.json
tools/normalize_experts_extra_config_key.py apply / --check / --revert the metadata fix
images/
ar_dit_full_2gpu.png AR+DiT, 2 GPUs (hero; AR branch ratio_idx=9)
dit_only_tp2.png this build, DiT-only TP2
bf16_dit_only_reference.png BF16 base, DiT-only TP2 (same session)
mxfp8_ar_dit_only.png -MXFP8 (auto_round), DiT-only TP2 (same session)
sibling_mxfp4_ar_dit_only.png -MXFP4-ar, DiT-only TP2 (same session)
grid_ditonly_fourway.png BF16 / MXFP8 / MXFP4-ar / this build + PSNR+SSIM labels
grid_ar_dit_branches.png three AR+DiT outputs, each labelled with its AR branch
ar_only_output.txt AR-only text (1168 B; see the AR-only caveat)
*.py / tokenizer* / utils/ inherited from the base model
Reproduce the analysis tables with
quantization_analysis/scripts/module_scheme_inventory.py <dir> --show-groups
(archived outputs: quantization_analysis/reference/scheme_templates/).
- Downloads last month
- 10






