Qwen3.8-Flash-Next REAP-288 (bf16)

Precision Disk What it's for
This build bf16 (native) 232 GB the source: requantize, fine-tune, max-quality serving
REAP-288 MLX 4-bit 4-bit 68 GB what most people should run

This is the full-precision REAP-288 checkpoint. It is the 288-of-512 expert build spliced directly from the original bf16 weights, at native precision, before any quantization. Every other REAP-288 artifact we publish, the MLX 4-bit, any GGUF or AWQ conversion, derives from these bytes. This checkpoint exists for the cases where you need the full-precision source: requantizing to a format of your choice, fine-tuning, or serving at maximum fidelity.

You do not need pmlx to use this model. The runnable REAP-288 build loads and serves on stock mlx-vlm today, with no patches and no extra engine. omlx v0.6.4+ also runs the MLX builds today and streams the n-gram table to a lower footprint. pmlx is a separate inference engine that is not public yet; it releases alongside the tiered model and makes the same MLX builds faster (see the table below), but everything you need to run REAP-288 today works without it.

If you just want to run the model

Use the 4-bit build: sh0wie/Qwen3.8-Flash-Next-REAP-288-MLX-4bit. It loads on stock mlx-vlm with no patches at ~68 GB resident, drops to ~39 GB when a streaming runtime keeps the n-gram table on NVMe (omlx v0.6.4+ does this today), decodes at ~28 tok/s on an M4 Max on stock mlx-vlm, and scores 91.5% on HumanEval. It is a third of this checkpoint's size and its quality is within a point of native precision. For most uses, that is the model you want.

pip install git+https://github.com/Blaizzy/mlx-vlm.git
python -m mlx_vlm.server \
  --model sh0wie/Qwen3.8-Flash-Next-REAP-288-MLX-4bit --port 8080

Decode speed for the 4-bit build

The ~28 tok/s stock number is what mlx-vlm gives you today with the n-gram table resident at ~68 GB. omlx streams the table to ~39 GB. The pmlx engine runs the same weights faster and ships with the tiered release.

Engine Resident Decode (M4 Max) Available
stock mlx-vlm, table resident 68 GB ~28 tok/s today
omlx, table streamed from NVMe 39 GB not separately benchmarked today (v0.6.4+)
pmlx, table streamed from NVMe 39 GB ~37 tok/s with the tiered release
pmlx, table in RAM 73 GB ~65 peak, ~41 sustained at 4K with the tiered release

Same 4-bit build in every row. The only measured public-runtime decode today is ~28 tok/s on stock mlx-vlm with the table resident; the 39 GB footprint comes from omlx streaming the table from NVMe, which we have not separately benchmarked. mlx-vlm streams the table once PR #2045 lands in a release, since the 4-bit build already ships the manifest it reads. The ~65 tok/s pmlx figure is the RAM-resident table and a short-context peak, not a sustained number.

About this build

  • 180B-parameter class: 125B main model, 51B n-gram embedding table, 48 layers alternating Gated DeltaNet and Qwen sparse attention, each with a 288-expert MoE routing top-10.
  • 288 of the original 512 experts per MoE layer, chosen by REAP saliency calibrated on the quantized weights over ~686K tokens of agentic-coding traffic. The kept-expert set is identical to the 4-bit build.
  • Native bf16. No quantization error is baked in.

What the bf16 source is for

  • Requantizing. This is the reference source. Convert it to MLX at any bit width, to GGUF for llama.cpp, to AWQ or an FP variant, without inheriting a prior quantization's rounding. Start from bf16 and you lose accuracy exactly once.
  • Fine-tuning. LoRA or full fine-tunes want full-precision weights. The 4-bit build is for inference; this is what you train on top of.
  • Max-quality serving. If you have the memory budget and want the highest fidelity the pruned model can produce, serve these weights directly. At 232 GB that means paging on a single Mac; the details are below.
Running the bf16 source on the pmlx engine (coming with the tiered release)

pmlx is a pure-MLX inference engine for Apple Silicon. It is not public yet; it releases alongside the tiered model, and the install URL below goes live then. None of this is needed to run the 4-bit build on mlx-vlm today. It is here because pmlx is what lets you serve the 232 GB bf16 source on a single Mac at all, by keeping part of the model resident and paging the rest from SSD.

# available when pmlx ships with the tiered release
pip install git+https://github.com/gethamster/pmlx
python -m pmlx.server \
  --model sh0wie/Qwen3.8-Flash-Next-REAP-288-bf16 --port 8123

The server is OpenAI-compatible (/v1/chat/completions, SSE streaming).

Residency is automatic. On load, pmlx reads how much memory your Mac has, reserves room for the KV cache (which grows with context length), and keeps as many experts resident on the GPU as fit; the rest page in from SSD as they are needed. On a 232 GB model that paging is how the model runs at all, and pmlx sets the resident fraction to whatever your machine can hold. The larger n-gram table streams from NVMe rather than sitting in RAM. Nothing is baked into the checkpoint, so the same download behaves correctly on differently-sized Macs without a re-download.

Pinning the ratio. When you want to set the resident fraction yourself, --resident-ratio does it:

python -m pmlx.server --model sh0wie/Qwen3.8-Flash-Next-REAP-288-bf16 \
  --resident-ratio 0.40 --port 8123

Resident memory follows one relation: a fixed backbone, plus room for the KV cache, plus the share of expert weights you keep on the GPU.

resident ~= backbone + KV_reserve + resident_ratio * expert_weight

The n-gram table is not in that sum; pmlx reads it from NVMe a few hundred bytes at a time, bit-identical to holding it in memory. So the only term you scale is the expert share. Keep 40% of the experts and you carry roughly 40% of the expert weight. That is what lets the full-precision model serve on a Mac: this 232 GB checkpoint serves in about 63 GB resident at a 0.40 ratio, with the remaining experts paging on demand and the table on SSD. A higher ratio keeps more experts hot and costs more memory; a lower ratio does the reverse. Quality does not move with the dial, because paged, cached, and prefetched expert gathers are bit-identical to the fully-resident path.

Serving bf16 pages far more bytes per token than the 4-bit build does, so throughput is lower and we do not quote a tok/s figure for it here. Treat this checkpoint as the quality-and-source tier, and the 4-bit build as the one you reach for when you want speed.

Provenance

  • Qwen/Qwen3.8-Flash-Next: upstream weights.
  • This build: REAP expert pruning 512 -> 288 per layer, calibrated on-device over ~686K tokens of agentic-coding traffic, then materialized at native bf16. The per-layer kept-expert manifest is the same selection shipped as reap_kept_experts.json in the 4-bit build.

Limitations

  • Calibration reflects one team's agentic-coding distribution. Retention numbers should not be read as general-domain; domains far from code may degrade more.
  • This is a source checkpoint, not a convenience build. At 232 GB it does not fit in RAM on a single Mac; serving it directly needs a paging engine such as pmlx (see above) or your own large-model stack. To simply run the model, take the 4-bit build, which loads on stock mlx-vlm.
  • Single-run evaluations on the lineage, no confidence intervals. Vision input is untested after pruning.

License

Qwen Community License 1.0, inherited from the base model Qwen/Qwen3.8-Flash-Next.

Downloads last month
916
Safetensors
Model size
125B params
Tensor type
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sh0wie/Qwen3.8-Flash-Next-REAP-288-bf16

Finetuned
(61)
this model
Quantizations
1 model

Collection including sh0wie/Qwen3.8-Flash-Next-REAP-288-bf16