Instructions to use sokann/Qwen3.6-27B-GGUF-4.256bpw with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf sokann/Qwen3.6-27B-GGUF-4.256bpw # Run inference directly in the terminal: llama cli -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf sokann/Qwen3.6-27B-GGUF-4.256bpw # Run inference directly in the terminal: llama cli -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf sokann/Qwen3.6-27B-GGUF-4.256bpw # Run inference directly in the terminal: ./llama-cli -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf sokann/Qwen3.6-27B-GGUF-4.256bpw # Run inference directly in the terminal: ./build/bin/llama-cli -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Use Docker
docker model run hf.co/sokann/Qwen3.6-27B-GGUF-4.256bpw
- LM Studio
- Jan
- Ollama
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with Ollama:
ollama run hf.co/sokann/Qwen3.6-27B-GGUF-4.256bpw
- Unsloth Desktop
- Pi
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sokann/Qwen3.6-27B-GGUF-4.256bpw" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with Docker Model Runner:
docker model run hf.co/sokann/Qwen3.6-27B-GGUF-4.256bpw
- Lemonade
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull sokann/Qwen3.6-27B-GGUF-4.256bpw
Run and chat with the model
lemonade run user.Qwen3.6-27B-GGUF-4.256bpw-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sokann/Qwen3.6-27B-GGUF-4.256bpw
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sokann/Qwen3.6-27B-GGUF-4.256bpw with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf sokann/Qwen3.6-27B-GGUF-4.256bpw
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sokann/Qwen3.6-27B-GGUF-4.256bpw" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.6-27B-GGUF-4.256bpw
This is a 4.256 BPW quantized model for the GPU poors with 16 GiB of VRAM. It works in both ik_llama.cpp and mainline llama.cpp.
It was quantized using the simplest ever recipe โ Q8_0 for the tiny ssm_alpha and ssm_beta tensors, IQ4_XS for the rest.
From local testing with llama-perplexity (wiki.test.raw, 580 chunks), it has the best quality and speed in the same size class:
| quant | this | bartowski Q3_K_M | unsloth UD-Q3_K_XL | mradermacher i1.IQ4_XS | bartowski IQ4_XS | unsloth IQ4_XS |
|---|---|---|---|---|---|---|
| Size (BPW) | 4.256 | 4.270 | 4.302 | 4.483 | 4.556 | 4.589 |
| Size (GiB) | 13.327 | 13.370 | 13.469 | 14.036 | 14.266 | 14.369 |
| VRAM usage (GiB) | 12.698 | 12.861 | 12.803 | 13.407 | 13.637 | 13.703 |
| Mean PPL(Q) | 7.098696 ยฑ 0.047344 | 6.993009 ยฑ 0.046208 | 6.995519 ยฑ 0.046227 | 7.020660 ยฑ 0.046587 | 6.996323 ยฑ 0.046332 | 6.950126 ยฑ 0.045846 |
| Mean PPL(base) | 6.908506 ยฑ 0.045543 | 6.908506 ยฑ 0.045543 | 6.908506 ยฑ 0.045543 | 6.908506 ยฑ 0.045543 | 6.908506 ยฑ 0.045543 | 6.908506 ยฑ 0.045543 |
| Cor(ln(PPL(Q)), ln(PPL(base))) | 99.19% | 98.52% | 98.82% | 99.30% | 99.32% | 99.38% |
| Mean KLD | 0.033452 ยฑ 0.000723 | 0.058818 ยฑ 0.000881 | 0.046348 ยฑ 0.000841 | 0.027289 ยฑ 0.000660 | 0.026270 ยฑ 0.000653 | 0.024728 ยฑ 0.000603 |
| Maximum KLD | 23.255085 | 24.616274 | 24.175169 | 18.568180 | 22.992002 | 21.687405 |
| 99.9% KLD | 2.907350 | 3.986622 | 3.614290 | 2.667850 | 2.385293 | 2.201674 |
| RMS ฮp | 4.936 ยฑ 0.054 % | 6.690 ยฑ 0.059 % | 5.867 ยฑ 0.060 % | 4.449 ยฑ 0.057 % | 4.352 ยฑ 0.057 % | 4.264 ยฑ 0.056 % |
| Same top p | 92.427 ยฑ 0.069 % | 90.350 ยฑ 0.077 % | 91.829 ยฑ 0.071 % | 93.903 ยฑ 0.062 % | 93.888 ยฑ 0.062 % | 93.997 ยฑ 0.062 % |
- Compared to Q3_K_M from bartowski and UD-Q3_K_XL from unsloth, this IQ4_XS quant uses slightly less VRAM while having better quality.
- The IQ4_XS quant from mradermacher, bartowski, and unsloth have much better quality, but they use more VRAM and are harder to fit into 16 GiB of VRAM.
With 16 GiB of VRAM, we can fit a context size of 65536 with quantized KV cache:
# mainline llama.cpp
-c 65536 -ctk q8_0 -ctv q8_0 -np 1
For brave souls that seek the TurboQuant experience (see #21038), we can also fit a context size of 128000 with more heavily quantized KV cache:
# mainline llama.cpp
-c 128000 -ctk q4_0 -ctv q4_0 -np 1
Size
Size from llama-server output:
llm_load_print_meta: model size = 13.327 GiB (4.256 BPW)
llm_load_print_meta: repeating layers = 12.069 GiB (4.257 BPW, 24.353 B parameters)
...
llm_load_tensors: CUDA_Host buffer size = 644.14 MiB
llm_load_tensors: CUDA0 buffer size = 13003.14 MiB
Recipe
blk\..*\.attn_q\.weight=iq4_xs
blk\..*\.attn_k\.weight=iq4_xs
blk\..*\.attn_v\.weight=iq4_xs
blk\..*\.attn_output\.weight=iq4_xs
blk\..*\.attn_gate\.weight=iq4_xs
blk\..*\.attn_qkv\.weight=iq4_xs
blk\..*\.ssm_alpha\.weight=q8_0
blk\..*\.ssm_beta\.weight=q8_0
blk\..*\.ssm_out\.weight=iq4_xs
blk\..*\.ffn_down\.weight=iq4_xs
blk\..*\.ffn_(gate|up)\.weight=iq4_xs
token_embd\.weight=iq4_xs
output\.weight=iq4_xs
Speed
llama-sweep-bench result with a RTX 3090, with flags -ngl 99 -mqkv -muge -cuda graphs=1 -c 128000 -wgt 1 -wb:
| PP | TG | N_KV | T_PP s | S_PP t/s | T_TG s | S_TG t/s |
|---|---|---|---|---|---|---|
| 512 | 128 | 0 | 0.335 | 1526.19 | 2.632 | 48.64 |
| 512 | 128 | 10240 | 0.376 | 1362.66 | 2.787 | 45.93 |
| 512 | 128 | 20480 | 0.416 | 1231.97 | 2.870 | 44.60 |
| 512 | 128 | 30720 | 0.457 | 1119.71 | 2.964 | 43.19 |
| 512 | 128 | 40960 | 0.500 | 1024.24 | 3.080 | 41.56 |
| 512 | 128 | 51200 | 0.545 | 940.27 | 3.183 | 40.21 |
| 512 | 128 | 61440 | 0.589 | 868.63 | 3.277 | 39.06 |
| 512 | 128 | 71680 | 0.630 | 812.78 | 3.378 | 37.89 |
| 512 | 128 | 81920 | 0.673 | 760.29 | 3.497 | 36.60 |
| 512 | 128 | 92160 | 0.716 | 715.36 | 3.605 | 35.51 |
| 512 | 128 | 102400 | 0.761 | 672.98 | 3.696 | 34.64 |
| 512 | 128 | 112640 | 0.802 | 638.68 | 3.798 | 33.70 |
| 512 | 128 | 122880 | 0.843 | 607.28 | 3.917 | 32.68 |
Performance
This quant uses the imatrix from mradermacher. It performs well enough in long reasoning tasks and agentic tasks.
- Downloads last month
- 460
We're not able to determine the quantization variants.
Model tree for sokann/Qwen3.6-27B-GGUF-4.256bpw
Base model
Qwen/Qwen3.6-27B