Mixed Precision GGUF layer quantization of Qwen3-32B by Qwen

Original model: https://huggingface.co/Qwen/Qwen3-32B

The hybrid quant employs different quantization levels on a per layer basis to increased flexibility of trading off performance vs file size. Less parameter bits are used at deep layers and more bits at cortex layers to simultaneously optimize quantized size and model performance. K quants are used in all the layers for faster CPU processing on partially offloaded models or GPU processing on older GPUs.

The layer quants are as follows (refreshed on 4/22/2026):

   [0 ,"Q4_K_M"],[1 ,"Q4_K_S"],[2 ,"Q3_K_L"],[3 ,"Q4_K_S"],[4 ,"Q3_K_L"],[5 ,"Q3_K_L"],[6 ,"Q3_K_L"],[7 ,"Q3_K_L"],
   [8 ,"Q3_K_L"],[9 ,"Q3_K_L"],[10,"Q3_K_L"],[11,"Q3_K_L"],[12,"Q3_K_L"],[13,"Q3_K_L"],[14,"Q3_K_L"],[15,"Q3_K_L"],
   [16,"Q3_K_L"],[17,"Q3_K_L"],[18,"Q3_K_L"],[19,"Q3_K_L"],[20,"Q3_K_L"],[21,"Q3_K_L"],[22,"Q3_K_L"],[23,"Q3_K_L"],
   [24,"Q4_K_S"],[25,"Q3_K_L"],[26,"Q4_K_S"],[27,"Q3_K_L"],[28,"Q4_K_S"],[29,"Q3_K_L"],[30,"Q4_K_S"],[31,"Q3_K_L"],
   [32,"Q4_K_S"],[33,"Q3_K_L"],[34,"Q4_K_S"],[35,"Q3_K_L"],[36,"Q4_K_S"],[37,"Q3_K_L"],[38,"Q4_K_S"],[39,"Q3_K_L"],
   [40,"Q4_K_S"],[41,"Q3_K_L"],[42,"Q4_K_S"],[43,"Q3_K_L"],[44,"Q4_K_S"],[45,"Q3_K_L"],[46,"Q4_K_S"],[47,"Q3_K_L"],
   [48,"Q4_K_S"],[49,"Q4_K_S"],[50,"Q4_K_S"],[51,"Q4_K_S"],[52,"Q4_K_S"],[53,"Q4_K_S"],[54,"Q4_K_S"],[55,"Q4_K_S"],
   [56,"Q4_K_M"],[57,"Q4_K_S"],[58,"Q4_K_M"],[59,"Q4_K_L"],[60,"Q5_K_S"],[61,"Q5_K_M"],[62,"Q5_K_L"],[63,"Q6_K_S"]
   ]'
   FLAGS="--token-embedding-type Q4_K --output-tensor-type Q6_K --layer-types-high"

These quants were select based on performance optimization over a set of curated test prompts.

Comparison:

Quant size PPL Comment
IQ4_XS 17.9e9 7.8 default embed and output
Q4_K_H 18.5e9 7.8 Q4_K embed Q6_K output

Usage:

This is a dense RL model. By default it will crank out a think block delimited by

THINK_START="<think>\n"
THINK_STOP="\n</think>\n\n"

To bypass thinking inject the think block delimiters following the assistant prompt template. The model is strong with think blocked bypassed but less accurate on harder prompts. The model exhibits strong "common sense", and exhibits correct reasoning on prompts smaller models mostly miss. It is a dense 32G parameter model which is not practical to run on CPU and must be offloaded into GPU by hook or crook to get usable gen rates.

This model can run fully offloaded on a 24G VRAM. This 24G VRAM can be cobbled together with 2x4070 over RPC. Q8 KV cache can be used to expand context. The model can be efficiently speculated with Qwen3 0.6B. Example configs and gen rates for 2x4070 (1 RPC) with optional Qwen3 0.6B speculation running with a custom downstream speculator with fixed draft block size ND on llama.cpp:

Spec ND QKV Context size gen rate Comment
0 F16 18k 21 tps No draft loaded
0 Q8_0 30k 21 tps ""
4 F16 10k 38 tps Qwen3 0.6B draft
4 Q8_0 12k 38 tps ""

Full evals for Q4_K_H quant (non refreshed quant) are available at https://huggingface.co/spaces/steampunque/benchlm

Download the file from below:

Link Type Size/e9 B Notes
Qwen3-32B.Q4_K_H.gguf Q4_K_H 18.5e9 B ~IQ4_XS size with much higher perf.

A discussion thread about the hybrid layer quant approach can be found here on the llama.cpp git repository:

https://github.com/ggml-org/llama.cpp/discussions/13040

Downloads last month
28
GGUF
Model size
33B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for steampunque/Qwen3-32B-MP-GGUF

Base model

Qwen/Qwen3-32B
Quantized
(178)
this model