Sakura — MiMo-V2.6-Flash-MOPD P160 (expert-pruned, SSD-streaming ready)

Sakura MiMo logo

Sakura MiMo P160 is an experimental expert-pruned GGUF of Xiaomi's MiMo-V2.6-Flash-MOPD: 160 of 256 routed experts per MoE layer, about 53.79 GiB. It is built to run on a single 64 GB AMD Strix Halo (Radeon 8060S, Vulkan) with ~35 GiB on the GPU and the remaining experts computed from system RAM — the MiMo SSD Streaming setup described below. Focus: German instructions, long planning prompts and coding. Independent community build, not an official Xiaomi release.

Part of the Sakura Micro line: compact, locally runnable derivatives of large models (lines: Sakura, Sakura Mini, Sakura Micro).

Provenance

  • Base model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD, the MOPD2 upgrade of MiMo-V2.6-Flash-RL (309B total / 15B active, 48 layers, 256 experts, top-8). Its routers are bit-identical to the RL checkpoint, so an expert selection measured on RL transfers exactly.
  • Built from Xiaomi's original weights: only the kept experts were downloaded (HTTP ranges), the MXFP4 experts were repacked losslessly, FP8 tensors converted exactly like llama.cpp's converter (checked byte-identical against the reference on the RL release).
  • Expert selection (per layer): protect the experts carrying 80 % of the full model's routing on its own German answers to long coding/planning prompts (on-policy measurement), then the decode-trace hot experts, then REAP saliency bands and routed-token share decide. Selections tuned only on foreign text scored better perplexity but looped in chat; measuring on the model's own outputs fixed that.
  • Quantization: llama.cpp with an importance matrix (Baekpica's, cut to the kept experts). Experts IQ2_XXS/IQ2_XS in layers 12–35, IQ2_XS/IQ2_S elsewhere, down projections of layers 44/46/47 IQ3_XXS; attention IQ4_XS, dense FFN Q6_K, embeddings Q8_0, output Q6_K.

Release artifact

MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf

  • Quantization: mixed low-bit, mostly IQ2_XS experts (details above); IQ2_XS in the file name is the dominant type
  • Size: 57,761,339,136 Bytes (53.79 GiB)
  • SHA-256: 90dbd32b5d41ec902da81d0c369a8649cb29c059d7d9624a0c2f49511fab3b1b
  • Parameters: ~195B total after pruning (from 309B), ~15B active per token (unchanged)
  • The bundled SHA256SUMS file can be used to verify the download.

Measured results

Build Size PPL German PPL code Chat checks
Sakura MiMo P160 (this release, MOPD) 53.8 GiB 6.02 5.80 short 7/8 · German stress 6/6 · long texts 4/4 · roadmaps 2/2 at temperature 1.0 · no loops
same selection and recipe, RL checkpoint 53.8 GiB 6.02 5.87 roadmap: 1/2 at temperature 0 (rumination loop), 2/2 at 0.6
144 experts, RL (cut from TrevorJS IQ2) 53.5 GiB 6.12 6.10 short 7/8 · long 4/4 · roadmaps 2/2 at temperature 0
full MiMo-V2.6-Flash-RL (TrevorJS IQ2, reference) 91.9 GiB 5.04 4.23 —

Decode speed of this release: 9–10.5 tok/s at 16k context and 5.6–7.7 tok/s at 32k, measured with the desktop in normal parallel use (details below).

Perplexity: held-out German and code text, 8 × 512 tokens, not used for selection or calibration (lower is better). Speed: AMD Ryzen AI Max+ 395 / Radeon 8060S, 64 GB, Vulkan, MiMo SSD Streaming setup below, measured with the desktop in normal parallel use (browser, chat apps; free RAM near zero) — see the speed table below for the context-size effect. Chat checks: short factual/arithmetic prompts, long free-form texts, six German mid-length answers and two long German roadmap/deployment prompts with fixed closing lines. These are small functional checks, not a claim of general quality versus the original model.

Why 160 experts

We built and tested several expert counts from the same MiMo-V2.6-Flash family on this machine (held-out PPL German / code; lower is better):

Experts kept per layer Size PPL German PPL code Chat
128 (hot-expert selection) 48.0 GiB 14.50 7.30 German broken (word salad, loops)
128 (German-protected) 48.0 GiB 5.97 6.99 German fine, but code loops and format errors
144 (on-policy) 53.5 GiB 6.12 6.10 stable
160 (on-policy, this release) 53.8 GiB 6.02 5.80 stable, no loops at temperature 1.0
256 (full model, reference) 91.9 GiB 5.04 4.23 — (2.6 tok/s here, mostly paged from NVMe)

The 128/144-expert rows and the full-model reference are MiMo-V2.6-Flash-RL builds. The same 160-expert selection on the RL checkpoint scored 6.02 / 5.87, so the step from 144 to 160 helps on its own; MOPD adds a little on code.

Smaller REAP/pruning levels cost noticeable quality — at 128 experts German or code breaks down. 160 experts give the best quality among the variants that still run at a usable speed on a 64 GB machine: the file is barely larger than the 144-expert build because the middle layers use slightly smaller quant types. Keeping more experts would push more weights off the GPU and cost speed. 160 is the sweet spot among the variants we tested for 64 GB; we did not build 176+.

Running it with stock llama.cpp

Any recent llama.cpp with mimo2 support loads this GGUF. With enough memory for the whole file on the GPU:

llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 32768 --temp 1.0 --top-p 0.95

Use temperature 1.0 / top-p 0.95 (Xiaomi's recommendation): at 0–0.6, long thinking tasks sometimes fell into rumination loops in our tests; at 1.0 they did not.

MiMo SSD Streaming — running it with less VRAM

The model is larger than a 32 GB GPU. MiMo SSD Streaming keeps the attention, the outlier-heavy early layers and the late layers on the GPU and lets the CPU compute the experts of a block of middle layers (14–33, ~20 GiB) from RAM, with the NVMe as backing store. Components (all in runtime/ of this repo):

Component What it does
-ot "…ffn_(gate|up|down)_exps\.weight=CPU" keeps the experts of chosen layers in host memory (mmap); fits the model into ~35 GiB GPU
Swap-MoE --expert-streaming (ek15072809/Swap-MoE) + our GPU guard patch demand-pages those experts from NVMe instead of loading them all up front
Op-offload on (no --no-op-offload) prompt processing copies the used experts to the GPU (fast when the experts are in RAM)
Expert lock (LLAMA_EXPERT_LOCK_LIST, our patch, optional) pins the most-used CPU-side experts in RAM (VirtualLock) so multi-token batches do not page-fault
UTF-8 sanitizing in chat parsing (our patch) pruned models sometimes emit a broken emoji byte sequence; stock llama-server then fails with HTTP 500 — the reply is kept, broken bytes become U+FFFD

Speed on this machine (decode, tok/s; desktop in parallel use, free RAM near zero):

Mode Context Lock Decode Prefill (4k-token prompt)
fast (default) 16k none 9–10.5 13–25
long 32k hottest 30 % (~6 GiB) 5.6–7.7 14–18
(24k, for reference) 24k none 6.6–7.4 32

All rows use the default -ub 512, which gives the best overall speed. A larger ubatch trades decode for prefill:

-ub Prefill (2k prompt) Decode
512 (default) 12.7 9.6
1024 22.3 5.9
2048 34–39 (4k prompt) ~5.3

When the CPU-side experts do not fit in RAM they are read from the NVMe once per ubatch, so larger ubatches speed up prefill; but their bigger GPU buffers spill into shared system memory and take RAM away from those experts, which nearly halves decode. A larger ubatch only pays off when the prompt is more than about twice as long as the answer (for example a one-off read of large files); for coding with long reasoning keep 512. -ub 4096 ran out of GPU memory on this machine.

The context size matters because the KV cache and GPU buffers spill into shared system memory; at 16k enough RAM stays free for the CPU-side experts. On a quiet machine with the RAM to spare, the same runtime reached 8.9 tok/s decode, ~120 tok/s prefill and 15/18/17 tok/s on 2/4/8-token batches with the expert lock (measured with our 144-expert build). n-gram speculation (--spec-type ngram-mod) helped the RL-based builds when returning edited files, but not this MOPD build (about 50 % acceptance), so it is off by default.

Example for 64 GB Strix Halo (runtime/start_server_example.bat, long as first argument for the 32k mode):

llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 16384 -fit off --expert-streaming ^
  -ot "blk\.(1[4-9]|2[0-9]|3[0-3])\.ffn_(gate|up|down)_exps\.weight=CPU" ^
  --temp 1.0 --top-p 0.95

Smaller GPUs: move more middle layers to the CPU by widening the -ot layer range (each middle layer holds ~1 GiB of experts; keep layers 0–11 and 36–47 on the GPU if you can, they are the most sensitive). Speed depends on whether the CPU-side experts fit in RAM: up to ~20 GiB in RAM decode at 6–10 tok/s here; with 56 GiB paged from NVMe (the full, unpruned model) it was 2.6 tok/s. The lock list (hotlock_L14-33.txt) only applies to layers 14–33; regenerate it with hot_lock_list.py for other ranges or leave the variable unset.

The runtime patches and build steps are in runtime/patches/ (llama.cpp f5e85d43a + Swap-MoE + ours). Stock llama.cpp ignores LLAMA_EXPERT_LOCK_LIST and runs the model without that feature.

Known limitations

  • Pruning removes capacity: expect less world knowledge than the original (e.g. a wrong statute number in a German legal answer) and occasional character-level typos in long German texts.
  • First-shot code is mediocre; review it.
  • Reversing a word letter by letter fails (all pruned variants).
  • Experimental release; speed figures are from one machine.

License

MIT, like the base model. The license text is provided in LICENSE; sources and credits in NOTICE.md.

Downloads last month
486
GGUF
Model size
195B params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF

Quantized
(23)
this model

Collection including webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF