Ornith-1.5-35B-A3B-oQ3e-mtp

This model was quantized using oQ (oMLX v0.6.2) mixed-precision quantization.

Base model: ornith-ai/Ornith-1.5-35B-A3B

Chat template: froggeric/Qwen-Fixed-Chat-Templates (v21.3, original backed up as chat_template.jinja.bak)

Quantization details

  • Model type: qwen3_5_moe
  • Bits: 3
  • Group size: 64
  • Mode: affine
  • Format: MLX safetensors
  • MTP: Preserved (mtp_num_hidden_layers: 1 — built-in MTP head, no external donor)
  • Calibration: oQ3e (enhanced, imatrix-based)
  • Note: text-only variant (vision tower dropped, −0.89 GB) to keep the MTP head intact

Environment

  • Hardware: M5 MacBook Air 32GB
  • Inference Framework: oMLX v0.6.2
  • Max Concurrent Requests: 4
  • Settings:
    • Thinking: Disabled
    • TurboQuant KV Cache: Enabled (4-bit)
    • Lightning MTP: Disabled (recommended — see note below)

⚠️ Lightning MTP

中文:本模型建议不要开启 Lightning MTP。MoE 架构模型 MTP 效果有差异,请先自行测试再决定是否开启。本卡的性能与智能基准数据均在不开启 Lightning MTP 的情况下测得。

English: For this model it is recommended to leave Lightning MTP disabled. MTP effectiveness varies across MoE-architecture models — test on your own setup before deciding whether to enable it. All benchmark results on this card were measured with Lightning MTP off.

Performance Benchmarks

Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.

Single Request Results

Test TTFT(ms) TPOT(ms) pp TPS tg TPS E2E(s) Throughput Peak Mem
pp1024/tg128 1130.5 19.60 905.8 tok/s 51.4 tok/s 3.633 317.1 tok/s 16.15 GB
pp4096/tg128 3824.1 20.13 1071.1 tok/s 50.1 tok/s 6.398 660.2 tok/s 16.88 GB

Continuous Batching (pp1024 / tg128)

Batch tg TPS Speedup pp TPS pp TPS/req TTFT(ms) E2E(s)
1x 51.4 tok/s 1.00x 905.8 tok/s 905.8 tok/s 1130.5 3.633
2x 72.0 tok/s 1.40x 631.6 tok/s 315.8 tok/s 2641.1 6.797
4x 114.7 tok/s 2.23x 517.6 tok/s 129.4 tok/s 4694.4 12.377

Intelligence Benchmark

Note: Each benchmark round tests only 30 questions. Results are for reference only.

Benchmark Accuracy Correct Total Time(s) Think
MMLU 56.7% 17 30 43.0 No
TRUTHFULQA 90.0% 27 30 24.5 No
GSM8K 93.3% 28 30 81.3 No
MATHQA 53.3% 16 30 50.6 No
HUMANEVAL 86.7% 26 30 98.0 No
Downloads last month
528
Safetensors
Model size
36B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including mlx-works/Ornith-1.5-35B-A3B-oQ3e-mtp