tiny-random-GlmMoeDsa-NVFP4

A tiny random model for testing, shrunk from nvidia/GLM-5.2-NVFP4: the same architecture, quantization config and checkpoint layout at test sizes. Its key patterns, dtypes and tensor ranks match the real checkpoint's (scripts/extract_layout.py).

modelopt NVFP4 on the routed experts only (the dense-first layer, attention, shared experts and MTP layer stay bf16), gate and up sharing one weight_scale_2, input_scale of 1.0 — as the real export ships it. Quantized with modelopt's own NVFP4QTensor.

reference/ holds the same weights dequantized to bf16, under the unquantized model's keys: the reference to compare logits against, so a test measures what the load path and kernels add, not the quantization itself.

The weights are random; the outputs mean nothing. scripts/ rebuilds it from the real checkpoint's config.json.

Downloads last month
618
Safetensors
Model size
10.8M params
Tensor type
F32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support