Instructions to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Use Docker
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- LM Studio
- Jan
- vLLM
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Ollama
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Unsloth Desktop
- Pi
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
- Lemonade
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Run and chat with the model
lemonade run user.Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF-IQ2_XS
List all available models
lemonade list
- Hermes Agent
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF:IQ2_XS" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sakura — MiMo-V2.6-Flash-MOPD P160 (expert-pruned, SSD-streaming ready)
Sakura MiMo P160 is an experimental expert-pruned GGUF of Xiaomi's MiMo-V2.6-Flash-MOPD: 160 of 256 routed experts per MoE layer, about 53.79 GiB. It is built to run on a single 64 GB AMD Strix Halo (Radeon 8060S, Vulkan) with ~35 GiB on the GPU and the remaining experts computed from system RAM — the MiMo SSD Streaming setup described below. Focus: German instructions, long planning prompts and coding. Independent community build, not an official Xiaomi release.
Part of the Sakura Micro line: compact, locally runnable derivatives of large models (lines: Sakura, Sakura Mini, Sakura Micro).
Provenance
- Base model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD, the MOPD2 upgrade of MiMo-V2.6-Flash-RL (309B total / 15B active, 48 layers, 256 experts, top-8). Its routers are bit-identical to the RL checkpoint, so an expert selection measured on RL transfers exactly.
- Built from Xiaomi's original weights: only the kept experts were downloaded (HTTP ranges), the MXFP4 experts were repacked losslessly, FP8 tensors converted exactly like llama.cpp's converter (checked byte-identical against the reference on the RL release).
- Expert selection (per layer): protect the experts carrying 80 % of the full model's routing on its own German answers to long coding/planning prompts (on-policy measurement), then the decode-trace hot experts, then REAP saliency bands and routed-token share decide. Selections tuned only on foreign text scored better perplexity but looped in chat; measuring on the model's own outputs fixed that.
- Quantization: llama.cpp with an importance matrix (Baekpica's, cut to the kept experts). Experts IQ2_XXS/IQ2_XS in layers 12–35, IQ2_XS/IQ2_S elsewhere, down projections of layers 44/46/47 IQ3_XXS; attention IQ4_XS, dense FFN Q6_K, embeddings Q8_0, output Q6_K.
Release artifact
MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf
- Quantization: mixed low-bit, mostly IQ2_XS experts (details above);
IQ2_XSin the file name is the dominant type - Size: 57,761,339,136 Bytes (53.79 GiB)
- SHA-256:
90dbd32b5d41ec902da81d0c369a8649cb29c059d7d9624a0c2f49511fab3b1b - Parameters: ~195B total after pruning (from 309B), ~15B active per token (unchanged)
- The bundled
SHA256SUMSfile can be used to verify the download.
Measured results
| Build | Size | PPL German | PPL code | Chat checks |
|---|---|---|---|---|
| Sakura MiMo P160 (this release, MOPD) | 53.8 GiB | 6.02 | 5.80 | short 7/8 · German stress 6/6 · long texts 4/4 · roadmaps 2/2 at temperature 1.0 · no loops |
| same selection and recipe, RL checkpoint | 53.8 GiB | 6.02 | 5.87 | roadmap: 1/2 at temperature 0 (rumination loop), 2/2 at 0.6 |
| 144 experts, RL (cut from TrevorJS IQ2) | 53.5 GiB | 6.12 | 6.10 | short 7/8 · long 4/4 · roadmaps 2/2 at temperature 0 |
| full MiMo-V2.6-Flash-RL (TrevorJS IQ2, reference) | 91.9 GiB | 5.04 | 4.23 | — |
Decode speed of this release: 9–10.5 tok/s at 16k context and 5.6–7.7 tok/s at 32k, measured with the desktop in normal parallel use (details below).
Perplexity: held-out German and code text, 8 × 512 tokens, not used for selection or calibration (lower is better). Speed: AMD Ryzen AI Max+ 395 / Radeon 8060S, 64 GB, Vulkan, MiMo SSD Streaming setup below, measured with the desktop in normal parallel use (browser, chat apps; free RAM near zero) — see the speed table below for the context-size effect. Chat checks: short factual/arithmetic prompts, long free-form texts, six German mid-length answers and two long German roadmap/deployment prompts with fixed closing lines. These are small functional checks, not a claim of general quality versus the original model.
Why 160 experts
We built and tested several expert counts from the same MiMo-V2.6-Flash family on this machine (held-out PPL German / code; lower is better):
| Experts kept per layer | Size | PPL German | PPL code | Chat |
|---|---|---|---|---|
| 128 (hot-expert selection) | 48.0 GiB | 14.50 | 7.30 | German broken (word salad, loops) |
| 128 (German-protected) | 48.0 GiB | 5.97 | 6.99 | German fine, but code loops and format errors |
| 144 (on-policy) | 53.5 GiB | 6.12 | 6.10 | stable |
| 160 (on-policy, this release) | 53.8 GiB | 6.02 | 5.80 | stable, no loops at temperature 1.0 |
| 256 (full model, reference) | 91.9 GiB | 5.04 | 4.23 | — (2.6 tok/s here, mostly paged from NVMe) |
The 128/144-expert rows and the full-model reference are MiMo-V2.6-Flash-RL builds. The same 160-expert selection on the RL checkpoint scored 6.02 / 5.87, so the step from 144 to 160 helps on its own; MOPD adds a little on code.
Smaller REAP/pruning levels cost noticeable quality — at 128 experts German or code breaks down. 160 experts give the best quality among the variants that still run at a usable speed on a 64 GB machine: the file is barely larger than the 144-expert build because the middle layers use slightly smaller quant types. Keeping more experts would push more weights off the GPU and cost speed. 160 is the sweet spot among the variants we tested for 64 GB; we did not build 176+.
Running it with stock llama.cpp
Any recent llama.cpp with mimo2 support loads this GGUF. With enough memory for the whole file on the GPU:
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 32768 --temp 1.0 --top-p 0.95
Use temperature 1.0 / top-p 0.95 (Xiaomi's recommendation): at 0–0.6, long thinking tasks sometimes fell into rumination loops in our tests; at 1.0 they did not.
MiMo SSD Streaming — running it with less VRAM
The model is larger than a 32 GB GPU. MiMo SSD Streaming keeps the attention, the outlier-heavy early layers and the
late layers on the GPU and lets the CPU compute the experts of a block of middle layers (14–33, ~20 GiB) from RAM,
with the NVMe as backing store. Components (all in runtime/ of this repo):
| Component | What it does |
|---|---|
-ot "…ffn_(gate|up|down)_exps\.weight=CPU" |
keeps the experts of chosen layers in host memory (mmap); fits the model into ~35 GiB GPU |
Swap-MoE --expert-streaming (ek15072809/Swap-MoE) + our GPU guard patch |
demand-pages those experts from NVMe instead of loading them all up front |
Op-offload on (no --no-op-offload) |
prompt processing copies the used experts to the GPU (fast when the experts are in RAM) |
Expert lock (LLAMA_EXPERT_LOCK_LIST, our patch, optional) |
pins the most-used CPU-side experts in RAM (VirtualLock) so multi-token batches do not page-fault |
| UTF-8 sanitizing in chat parsing (our patch) | pruned models sometimes emit a broken emoji byte sequence; stock llama-server then fails with HTTP 500 — the reply is kept, broken bytes become U+FFFD |
Speed on this machine (decode, tok/s; desktop in parallel use, free RAM near zero):
| Mode | Context | Lock | Decode | Prefill (4k-token prompt) |
|---|---|---|---|---|
| fast (default) | 16k | none | 9–10.5 | 13–25 |
| long | 32k | hottest 30 % (~6 GiB) | 5.6–7.7 | 14–18 |
| (24k, for reference) | 24k | none | 6.6–7.4 | 32 |
All rows use the default -ub 512, which gives the best overall speed. A larger ubatch trades decode for prefill:
-ub |
Prefill (2k prompt) | Decode |
|---|---|---|
| 512 (default) | 12.7 | 9.6 |
| 1024 | 22.3 | 5.9 |
| 2048 | 34–39 (4k prompt) | ~5.3 |
When the CPU-side experts do not fit in RAM they are read from the NVMe once per ubatch, so larger ubatches speed up
prefill; but their bigger GPU buffers spill into shared system memory and take RAM away from those experts, which
nearly halves decode. A larger ubatch only pays off when the prompt is more than about twice as long as the answer
(for example a one-off read of large files); for coding with long reasoning keep 512. -ub 4096 ran out of GPU
memory on this machine.
The context size matters because the KV cache and GPU buffers spill into shared system memory; at 16k enough RAM
stays free for the CPU-side experts. On a quiet machine with the RAM to spare, the same runtime reached 8.9 tok/s
decode, ~120 tok/s prefill and 15/18/17 tok/s on 2/4/8-token batches with the expert lock (measured with our
144-expert build). n-gram speculation (--spec-type ngram-mod) helped the RL-based builds when returning edited files,
but not this MOPD build (about 50 % acceptance), so it is off by default.
Example for 64 GB Strix Halo (runtime/start_server_example.bat, long as first argument for the 32k mode):
llama-server -m MiMo-V2.6-Flash-MOPD-P160-OnPolicy-v3la-IQ2_XS.gguf -ngl 99 --jinja -c 16384 -fit off --expert-streaming ^
-ot "blk\.(1[4-9]|2[0-9]|3[0-3])\.ffn_(gate|up|down)_exps\.weight=CPU" ^
--temp 1.0 --top-p 0.95
Smaller GPUs: move more middle layers to the CPU by widening the -ot layer range (each middle layer holds
~1 GiB of experts; keep layers 0–11 and 36–47 on the GPU if you can, they are the most sensitive). Speed depends on
whether the CPU-side experts fit in RAM: up to ~20 GiB in RAM decode at 6–10 tok/s here; with 56 GiB paged from NVMe
(the full, unpruned model) it was 2.6 tok/s. The lock list (hotlock_L14-33.txt) only applies to layers 14–33;
regenerate it with hot_lock_list.py for other ranges or leave the variable unset.
The runtime patches and build steps are in runtime/patches/ (llama.cpp f5e85d43a + Swap-MoE + ours). Stock
llama.cpp ignores LLAMA_EXPERT_LOCK_LIST and runs the model without that feature.
Known limitations
- Pruning removes capacity: expect less world knowledge than the original (e.g. a wrong statute number in a German legal answer) and occasional character-level typos in long German texts.
- First-shot code is mediocre; review it.
- Reversing a word letter by letter fails (all pruned variants).
- Experimental release; speed figures are from one machine.
License
MIT, like the base model. The license text is provided in LICENSE; sources and credits in NOTICE.md.
- Downloads last month
- 486
2-bit
Model tree for webmp3/Sakura-MiMo-V2.6-Flash-MOPD-P160-GGUF
Base model
XiaomiMiMo/MiMo-V2.6-Flash-MOPD