Instructions to use mlx-works/Ornith-1.5-35B-A3B-oQ3e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-works/Ornith-1.5-35B-A3B-oQ3e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Ornith-1.5-35B-A3B-oQ3e-mtp mlx-works/Ornith-1.5-35B-A3B-oQ3e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Ornith-1.5-35B-A3B-oQ3e-mtp
This model was quantized using oQ (oMLX v0.6.2) mixed-precision quantization.
Base model: ornith-ai/Ornith-1.5-35B-A3B
Chat template: froggeric/Qwen-Fixed-Chat-Templates (v21.3, original backed up as chat_template.jinja.bak)
Quantization details
- Model type: qwen3_5_moe
- Bits: 3
- Group size: 64
- Mode: affine
- Format: MLX safetensors
- MTP: Preserved (mtp_num_hidden_layers: 1 — built-in MTP head, no external donor)
- Calibration: oQ3e (enhanced, imatrix-based)
- Note: text-only variant (vision tower dropped, −0.89 GB) to keep the MTP head intact
Environment
- Hardware: M5 MacBook Air 32GB
- Inference Framework: oMLX v0.6.2
- Max Concurrent Requests: 4
- Settings:
- Thinking: Disabled
- TurboQuant KV Cache: Enabled (4-bit)
- Lightning MTP: Disabled (recommended — see note below)
⚠️ Lightning MTP
中文:本模型建议不要开启 Lightning MTP。MoE 架构模型 MTP 效果有差异,请先自行测试再决定是否开启。本卡的性能与智能基准数据均在不开启 Lightning MTP 的情况下测得。
English: For this model it is recommended to leave Lightning MTP disabled. MTP effectiveness varies across MoE-architecture models — test on your own setup before deciding whether to enable it. All benchmark results on this card were measured with Lightning MTP off.
Performance Benchmarks
Note: Results are for reference only and may vary depending on hardware, software configuration, and workload.
Single Request Results
| Test | TTFT(ms) | TPOT(ms) | pp TPS | tg TPS | E2E(s) | Throughput | Peak Mem |
|---|---|---|---|---|---|---|---|
| pp1024/tg128 | 1130.5 | 19.60 | 905.8 tok/s | 51.4 tok/s | 3.633 | 317.1 tok/s | 16.15 GB |
| pp4096/tg128 | 3824.1 | 20.13 | 1071.1 tok/s | 50.1 tok/s | 6.398 | 660.2 tok/s | 16.88 GB |
Continuous Batching (pp1024 / tg128)
| Batch | tg TPS | Speedup | pp TPS | pp TPS/req | TTFT(ms) | E2E(s) |
|---|---|---|---|---|---|---|
| 1x | 51.4 tok/s | 1.00x | 905.8 tok/s | 905.8 tok/s | 1130.5 | 3.633 |
| 2x | 72.0 tok/s | 1.40x | 631.6 tok/s | 315.8 tok/s | 2641.1 | 6.797 |
| 4x | 114.7 tok/s | 2.23x | 517.6 tok/s | 129.4 tok/s | 4694.4 | 12.377 |
Intelligence Benchmark
Note: Each benchmark round tests only 30 questions. Results are for reference only.
| Benchmark | Accuracy | Correct | Total | Time(s) | Think |
|---|---|---|---|---|---|
| MMLU | 56.7% | 17 | 30 | 43.0 | No |
| TRUTHFULQA | 90.0% | 27 | 30 | 24.5 | No |
| GSM8K | 93.3% | 28 | 30 | 81.3 | No |
| MATHQA | 53.3% | 16 | 30 | 50.6 | No |
| HUMANEVAL | 86.7% | 26 | 30 | 98.0 | No |
- Downloads last month
- 528
3-bit