How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
# Run inference directly in the terminal:
llama cli -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
# Run inference directly in the terminal:
llama cli -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
# Run inference directly in the terminal:
./llama-cli -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
Use Docker
docker model run hf.co/pcuenq/MiMo-V2.6-Pro-MOPD-GGUF:
Quick Links

MiMo-V2.6-Pro-MOPD

Run with https://llama.app

llama serve -hf pcuenq/MiMo-V2.6-Pro-MOPD-GGUF

Source models

Quants

  • MXFP4
  • Q3_K
  • Q2_K

Notes

  • The MXFP4 quant keeps the routed experts at their native MXFP4 precision.
  • The Q3_K quant keeps the expert down projections at MXFP4, and quantizes the gate/up projections to Q3_K.
  • The Q2_K quant keeps the expert down projections at MXFP4, and quantizes the gate/up projections to Q2_K.
  • Includes MTP sidecars (Q4_0 and Q8_0) for speculative decoding (--mtp).
  • Includes a DFlash drafter sidecar (BF16 and Q8_0) for speculative decoding, converted from the dflash/ subdirectory of the source repo.
  • Includes a mmproj adapter (BF16 and Q8_0) for the vision and audio encoders.
  • The Q3_K and Q2_K expert gate/up tensors are calibrated with the imatrix from https://huggingface.co/AesSedai/MiMo-V2.6-Pro-MOPD-GGUF

This model was manually converted by pcuenq following the approach used in ggml-org/MiMo-V2.6-Flash-MOPD-GGUF.

Downloads last month
889
GGUF
Model size
1T params
Architecture
mimo2
Hardware compatibility
Log In to add your hardware

2-bit

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for pcuenq/MiMo-V2.6-Pro-MOPD-GGUF

Quantized
(7)
this model