Swift-Qwen3.8-27B-Splash

Change coming in the next Splash release. Splash's main now loads upstream MLX 4-bit (affine, group 64) and GGUF checkpoints directly and prepares its own weights (incoai/splash#133, #134), so pre-converted packages like this one become Splash's legacy path. This package keeps working, and on Splash 1.0.x (current) it is still the way to run this model. After the next release you should be able to skip it:

splash serve --model ukisai/Swift-Qwen3.8-27B-GGUF:Q4_K_M

This is the command Splash's maintainers give for Swift (#87); :Q6_K and :Q8_0 also load.

ukisai/Swift-Qwen3.8-27b packaged for Splash, Inco AI's inference engine for Apple silicon.

Newer release: Swift-1.5 is packaged for Splash at SiliconSpecies/Swift-1.5-4bit-MLX-Splash.

Splash documents the package shape but ships no converter, so the byte-level layout β€” section offsets, packing order, the quantization rule β€” was worked out from the engine source and verified against the one existing reference package. Independent and unofficial; not endorsed by UkisAI or Inco AI.

17.4 GB, uniform 4-bit (group 64, affine). Apple silicon + Splash only β€” which includes LM Studio's Splash runtime (LM Studio Bionic 1.1.5 or later: Settings β†’ Runtime β†’ Splash, then download this repository). It does not load in Transformers, MLX or llama.cpp.

Quick start

  1. Install Splash (Apple silicon, via Homebrew):

    brew install incoai/tap/splash
    
  2. Serve this model. The first run downloads the 17.4 GB package from Hugging Face:

    splash serve --model SiliconSpecies/Swift-Qwen3.8-27B-Splash
    

    Wait for Ready Β· SiliconSpecies/Swift-Qwen3.8-27B-Splash Β· … Β· http://127.0.0.1:8000. Leave it running; Ctrl+C stops it.

  3. Chat in the browser at http://127.0.0.1:8000.

  4. Or call the OpenAI-compatible API from another terminal. The model id is the repository name:

    curl http://127.0.0.1:8000/v1/chat/completions \
      -H 'Content-Type: application/json' \
      -d '{
        "model": "SiliconSpecies/Swift-Qwen3.8-27B-Splash",
        "messages": [{"role": "user", "content": "Write a haiku about the sea."}],
        "temperature": 1.0, "top_p": 0.95, "top_k": 20,
        "max_tokens": 2048
      }'
    

    Any OpenAI client works the same way: base URL http://127.0.0.1:8000/v1, model SiliconSpecies/Swift-Qwen3.8-27B-Splash, any API key unless you started the server with --api-key.

  5. Or connect a coding agent to the running server: splash claude, splash codex, splash opencode or splash hermes.

Send the sampling values in step 4 yourself β€” the API is greedy when they are omitted (see Sampling) β€” and leave room in max_tokens for thinking (see Reasoning effort). Useful splash serve options: --max-memory 28G, --max-context 100K, --api-key KEY, --no-webui.

part source treatment
target/ β€” 64 layers, embedding, head Swift BF16 quantized to 4-bit β€” this is the conversion
draft/ β€” DFlash 2 incoai/Qwen3.8-27B-Splash byte-identical copy
vision/ incoai/Qwen3.8-27B-Splash byte-identical copy (equals Qwen3.8-27B's own vision tower, 333/333 tensors)
tokenizer/ Swift unmodified, except one chat-template line patched as in the official package β€” see below

Quality

95 short-answer tasks with checkable answers (multi-step arithmetic, word problems, sequences, logic, code output, recall, units), greedy, identical scoring throughout:

configuration precision score
this package uniform 4-bit 95/95
incoai/Qwen3.8-27B-Splash (base model, same engine) uniform 4-bit 95/95
Swift via oMLX oQ4e mixed 4/5-bit 95/95
Swift via mlx_lm.convert uniform 4-bit 95/95

Indistinguishable from the official package and from higher-precision Swift builds at this difficulty. The set saturates, so it cannot rank them, and it is not a claim of parity with full-precision Swift β€” it cannot detect a small regression; the fidelity measurement below can. A standard harness run (GSM8K / ARC / MMLU) is pending.

Fidelity to BF16

Next-token KL divergence against BF16 Swift, on the corpora and protocol of the agentionai/Qwen3.8-27B-AP-GGUF card (60 Γ— 2,048 tokens per corpus, second half of each chunk scored):

KL, neutral web KL, wikitext-2 KL p99 (web / wiki) top-1 agreement perplexity
this package 0.032 0.047 0.23 / 0.43 90% +2.6% / +1.9%

For scale, that card's 4-bit GGUFs of base Qwen3.8-27B score 0.010–0.018 (web) and 0.015–0.024 (wikitext); this package sits nearer their 3-bit files (0.025 / 0.034). The cost is the format's: splash-packed-q4 stores every tensor, output head included, at uniform 4-bit, where llama.cpp's 4-bit mixes keep sensitive tensors at 5–6 bits. About one token in ten differs from BF16's top choice β€” invisible on the saturated set above, but real on long or hard generations.

Rounding is not what limits it: MLX's affine rule scores the same (0.032 / 0.048), and a clip-search variant that minimises weight error is no better (0.034 / 0.046). Activation-aware rounding within the same format is the remaining lever, untested so far.

Caveats: the GGUF figures are for base Qwen3.8 via llama.cpp, ours for Swift via MLX using this package's exact 4-bit values (verified bit-exact against the encoder), with a 16-bit KV cache β€” Splash's 8-bit KV cache adds a little more. KL is computed over the reference's top 64 tokens plus one tail bucket, which slightly understates it.

Speed

Measured on an M5 Max (40-core GPU, 128 GB) against the same weights in oMLX β€” whose oQ4e-mtp build is mixed 4/5-bit, i.e. slightly higher precision than this one.

Throughput against context length

prompt tokens prefill: this oMLX decode: this oMLX
523 659 510 123.6 90.2
2,059 824 718 129.2 71.7
8,203 867 850 87.8 64.0
32,779 772 785 93.2 55.9
65,547 703 642 87.8 36.6

Prefill is a wash. Decode is the difference and it widens with context β€” this package holds 79–129 tok/s across the range while oMLX decays from 90 to 37. At 64k that is 2.4Γ—.

DFlash 2 acceptance is 0.42–0.76 by prompt (mean ~0.61) against 0.501 for the base model the drafter was trained on. Decode rate tracks acceptance, which is most of the variance above. The drafter is reused unmodified; a Swift-targeted DFlash 2 would gain roughly a further 5%.

Measurement conditions

temperature: 0 sent explicitly to both (reproducibility, not Swift's recommended setting); max_tokens: 400; streamed, TTFT is wall time to the first token; tokens counted for every engine with this model's own tokenizer, because Splash streams one token per SSE delta while oMLX batches several; every prompt uniquely salted β€” without that a doubling sweep makes each prompt a prefix of the next and Splash's prefix cache serves half of it (a 32,779-token prompt logged cached 16,416 and reported 1,462 tok/s where the cold rate is 772); both engines otherwise at defaults, run back to back on an idle machine.

Sampling

Swift recommends temperature 1.0, top_k 20, top_p 0.95. The format has nowhere to carry generation_config.json β€” the installer accepts exactly five tokenizer files β€” and Splash deliberately keeps sampling out of the model.

  • Splash's chat page already sends those values. Browser users need do nothing.
  • The /v1 API defaults to greedy (temperature 0.0, top_p 1.0, top_k 0) when a client omits them. Send them explicitly.
  • If your client cannot β€” DSH, for one, has no top_p/top_k in its request model at all β€” a ~50-line proxy that fills them in works and costs no measurable latency.

Greedy is arguably the better default for agent and coding work. The failure mode is running it without realising.

Reasoning effort

Swift's chat template implements low, medium, xhigh, defaulting to xhigh. Splash passes reasoning_effort straight into the template, so a value outside that set raises a template error (high/max alias to xhigh, minimal to low).

reasoning_effort reasoning tokens total output
low 88 123
medium 83 163
xhigh 51 77
omitted 51 77 β€” confirming the default

Measured on one word problem, all answered correctly. The levels change the instructions the template writes, not a token budget, so higher effort can be shorter.

Watch out: at xhigh with small max_tokens, replies return empty content and finish_reason: "length" because the budget went to the thinking channel. Allow ~400+ output tokens, or use low for simple tool calls.

Conversion notes

Uniform 4-bit, group 64, affine (w = codeΒ·scale + bias, bf16 scale and bias), matching the rule the reference package uses β€” derived by decoding it and confirmed bit-exact against the base checkpoint. Per-section fidelity against the BF16 source: cosine 0.99501–0.99585, relative RMSE 0.091–0.100, across all 320 quantized sections.

If you attempt this yourself: every bf16 norm section stores Ξ³ + 1, including query-norm and key-norm, with gdn-norm the lone exception. Getting the attention norms wrong leaves the model fluent while destroying multi-step reasoning β€” and per-section error metrics will not show it.

Chat template. Swift's template rejects a system message anywhere but first, so clients that send a second one β€” JetBrains' commit-message generation through LM Studio's Splash runtime, for one β€” failed with messages could not be rendered. The official incoai/Qwen3.8-27B-Splash template renders such a message in place; this package applies the same one-line patch, and drops the unpatched copy from tokenizer_config.json as the official package does. Conversations that rendered before render byte-for-byte the same.

Licence

Swift contribution under the Swift Open License v1.0 (LICENSE); base model Apache 2.0 (LICENSE-APACHE-2.0). See NOTICE for attribution and the full list of changes.

Commercial use is free under US$1,000,000 gross annual revenue. Above that it requires a Swift Enterprise License from UkisAI. That condition travels with these weights and applies to you.

Credit: Swift fine-tune by UkisAI; Qwen3.8-27B by Alibaba Cloud; Splash engine, packed format and DFlash 2 drafter by Inco AI.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for SiliconSpecies/Swift-Qwen3.8-27B-Splash

Base model

Qwen/Qwen3.8-27B
Quantized
(59)
this model