Swift-Qwen3.8-27B-Splash
Change coming in the next Splash release. Splash's
mainnow loads upstream MLX 4-bit (affine, group 64) and GGUF checkpoints directly and prepares its own weights (incoai/splash#133, #134), so pre-converted packages like this one become Splash's legacy path. This package keeps working, and on Splash 1.0.x (current) it is still the way to run this model. After the next release you should be able to skip it:splash serve --model ukisai/Swift-Qwen3.8-27B-GGUF:Q4_K_MThis is the command Splash's maintainers give for Swift (#87);
:Q6_Kand:Q8_0also load.
ukisai/Swift-Qwen3.8-27b packaged for Splash, Inco AI's inference engine for Apple silicon.
Newer release: Swift-1.5 is packaged for Splash at
SiliconSpecies/Swift-1.5-4bit-MLX-Splash.
Splash documents the package shape but ships no converter, so the byte-level layout β section offsets, packing order, the quantization rule β was worked out from the engine source and verified against the one existing reference package. Independent and unofficial; not endorsed by UkisAI or Inco AI.
17.4 GB, uniform 4-bit (group 64, affine). Apple silicon + Splash only β which includes LM Studio's Splash runtime (LM Studio Bionic 1.1.5 or later: Settings β Runtime β Splash, then download this repository). It does not load in Transformers, MLX or llama.cpp.
Quick start
Install Splash (Apple silicon, via Homebrew):
brew install incoai/tap/splashServe this model. The first run downloads the 17.4 GB package from Hugging Face:
splash serve --model SiliconSpecies/Swift-Qwen3.8-27B-SplashWait for
Ready Β· SiliconSpecies/Swift-Qwen3.8-27B-Splash Β· β¦ Β· http://127.0.0.1:8000. Leave it running; Ctrl+C stops it.Chat in the browser at http://127.0.0.1:8000.
Or call the OpenAI-compatible API from another terminal. The model id is the repository name:
curl http://127.0.0.1:8000/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{ "model": "SiliconSpecies/Swift-Qwen3.8-27B-Splash", "messages": [{"role": "user", "content": "Write a haiku about the sea."}], "temperature": 1.0, "top_p": 0.95, "top_k": 20, "max_tokens": 2048 }'Any OpenAI client works the same way: base URL
http://127.0.0.1:8000/v1, modelSiliconSpecies/Swift-Qwen3.8-27B-Splash, any API key unless you started the server with--api-key.Or connect a coding agent to the running server:
splash claude,splash codex,splash opencodeorsplash hermes.
Send the sampling values in step 4 yourself β the API is greedy when they are omitted (see
Sampling) β and leave room in max_tokens for thinking (see
Reasoning effort). Useful splash serve options: --max-memory 28G,
--max-context 100K, --api-key KEY, --no-webui.
| part | source | treatment |
|---|---|---|
target/ β 64 layers, embedding, head |
Swift BF16 | quantized to 4-bit β this is the conversion |
draft/ β DFlash 2 |
incoai/Qwen3.8-27B-Splash |
byte-identical copy |
vision/ |
incoai/Qwen3.8-27B-Splash |
byte-identical copy (equals Qwen3.8-27B's own vision tower, 333/333 tensors) |
tokenizer/ |
Swift | unmodified, except one chat-template line patched as in the official package β see below |
Quality
95 short-answer tasks with checkable answers (multi-step arithmetic, word problems, sequences, logic, code output, recall, units), greedy, identical scoring throughout:
| configuration | precision | score |
|---|---|---|
| this package | uniform 4-bit | 95/95 |
incoai/Qwen3.8-27B-Splash (base model, same engine) |
uniform 4-bit | 95/95 |
Swift via oMLX oQ4e |
mixed 4/5-bit | 95/95 |
Swift via mlx_lm.convert |
uniform 4-bit | 95/95 |
Indistinguishable from the official package and from higher-precision Swift builds at this difficulty. The set saturates, so it cannot rank them, and it is not a claim of parity with full-precision Swift β it cannot detect a small regression; the fidelity measurement below can. A standard harness run (GSM8K / ARC / MMLU) is pending.
Fidelity to BF16
Next-token KL divergence against BF16 Swift, on the corpora and protocol of the
agentionai/Qwen3.8-27B-AP-GGUF
card (60 Γ 2,048 tokens per corpus, second half of each chunk scored):
| KL, neutral web | KL, wikitext-2 | KL p99 (web / wiki) | top-1 agreement | perplexity | |
|---|---|---|---|---|---|
| this package | 0.032 | 0.047 | 0.23 / 0.43 | 90% | +2.6% / +1.9% |
For scale, that card's 4-bit GGUFs of base Qwen3.8-27B score 0.010β0.018 (web) and
0.015β0.024 (wikitext); this package sits nearer their 3-bit files (0.025 / 0.034).
The cost is the format's: splash-packed-q4 stores every tensor, output head included, at
uniform 4-bit, where llama.cpp's 4-bit mixes keep sensitive tensors at 5β6 bits. About one
token in ten differs from BF16's top choice β invisible on the saturated set above, but real
on long or hard generations.
Rounding is not what limits it: MLX's affine rule scores the same (0.032 / 0.048), and a clip-search variant that minimises weight error is no better (0.034 / 0.046). Activation-aware rounding within the same format is the remaining lever, untested so far.
Caveats: the GGUF figures are for base Qwen3.8 via llama.cpp, ours for Swift via MLX using this package's exact 4-bit values (verified bit-exact against the encoder), with a 16-bit KV cache β Splash's 8-bit KV cache adds a little more. KL is computed over the reference's top 64 tokens plus one tail bucket, which slightly understates it.
Speed
Measured on an M5 Max (40-core GPU, 128 GB) against the same weights in oMLX β whose
oQ4e-mtp build is mixed 4/5-bit, i.e. slightly higher precision than this one.
| prompt tokens | prefill: this | oMLX | decode: this | oMLX |
|---|---|---|---|---|
| 523 | 659 | 510 | 123.6 | 90.2 |
| 2,059 | 824 | 718 | 129.2 | 71.7 |
| 8,203 | 867 | 850 | 87.8 | 64.0 |
| 32,779 | 772 | 785 | 93.2 | 55.9 |
| 65,547 | 703 | 642 | 87.8 | 36.6 |
Prefill is a wash. Decode is the difference and it widens with context β this package holds 79β129 tok/s across the range while oMLX decays from 90 to 37. At 64k that is 2.4Γ.
DFlash 2 acceptance is 0.42β0.76 by prompt (mean ~0.61) against 0.501 for the base model the drafter was trained on. Decode rate tracks acceptance, which is most of the variance above. The drafter is reused unmodified; a Swift-targeted DFlash 2 would gain roughly a further 5%.
Measurement conditions
temperature: 0 sent explicitly to both (reproducibility, not Swift's recommended
setting); max_tokens: 400; streamed, TTFT is wall time to the first token; tokens counted
for every engine with this model's own tokenizer, because Splash streams one token per
SSE delta while oMLX batches several; every prompt uniquely salted β without that a
doubling sweep makes each prompt a prefix of the next and Splash's prefix cache serves half
of it (a 32,779-token prompt logged cached 16,416 and reported 1,462 tok/s where the cold
rate is 772); both engines otherwise at
defaults, run back to back on an idle machine.
Sampling
Swift recommends temperature 1.0, top_k 20, top_p 0.95. The format has nowhere to carry
generation_config.json β the installer accepts exactly five tokenizer files β and Splash
deliberately keeps sampling out of the model.
- Splash's chat page already sends those values. Browser users need do nothing.
- The
/v1API defaults to greedy (temperature 0.0, top_p 1.0, top_k 0) when a client omits them. Send them explicitly. - If your client cannot β DSH, for one, has no
top_p/top_kin its request model at all β a ~50-line proxy that fills them in works and costs no measurable latency.
Greedy is arguably the better default for agent and coding work. The failure mode is running it without realising.
Reasoning effort
Swift's chat template implements low, medium, xhigh, defaulting to xhigh.
Splash passes reasoning_effort straight into the template, so a value outside that set
raises a template error (high/max alias to xhigh, minimal to low).
reasoning_effort |
reasoning tokens | total output |
|---|---|---|
low |
88 | 123 |
medium |
83 | 163 |
xhigh |
51 | 77 |
| omitted | 51 | 77 β confirming the default |
Measured on one word problem, all answered correctly. The levels change the instructions the template writes, not a token budget, so higher effort can be shorter.
Watch out: at xhigh with small max_tokens, replies return empty content and
finish_reason: "length" because the budget went to the thinking channel. Allow ~400+
output tokens, or use low for simple tool calls.
Conversion notes
Uniform 4-bit, group 64, affine (w = codeΒ·scale + bias, bf16 scale and bias), matching the
rule the reference package uses β derived by decoding it and confirmed bit-exact against the
base checkpoint. Per-section fidelity against the BF16 source: cosine 0.99501β0.99585,
relative RMSE 0.091β0.100, across all 320 quantized sections.
If you attempt this yourself: every bf16 norm section stores Ξ³ + 1, including
query-norm and key-norm, with gdn-norm the lone exception. Getting the attention norms
wrong leaves the model fluent while destroying multi-step reasoning β and per-section error
metrics will not show it.
Chat template. Swift's template rejects a system message anywhere but first, so clients
that send a second one β JetBrains' commit-message generation through LM Studio's Splash
runtime, for one β failed with messages could not be rendered. The official
incoai/Qwen3.8-27B-Splash template renders such a message in place; this package applies
the same one-line patch, and drops the unpatched copy from tokenizer_config.json as the
official package does. Conversations that rendered before render byte-for-byte the same.
Licence
Swift contribution under the Swift Open License v1.0 (LICENSE); base model Apache 2.0
(LICENSE-APACHE-2.0). See NOTICE for attribution and the full list of changes.
Commercial use is free under US$1,000,000 gross annual revenue. Above that it requires a Swift Enterprise License from UkisAI. That condition travels with these weights and applies to you.
Credit: Swift fine-tune by UkisAI; Qwen3.8-27B by Alibaba Cloud; Splash engine, packed format and DFlash 2 drafter by Inco AI.