Text Generation
GGUF
llama.cpp
quantized
conversational
imatrix
archsloth
autoround
kl-divergence
multilingual
korean
code
local-llm
cpu
on-device
qwen3
4b
Instructions to use Archsloth/Qwen3-4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Archsloth/Qwen3-4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Archsloth/Qwen3-4B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Archsloth/Qwen3-4B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Archsloth/Qwen3-4B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Archsloth/Qwen3-4B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Archsloth/Qwen3-4B-GGUF:Q4_K_M
- Ollama
How to use Archsloth/Qwen3-4B-GGUF with Ollama:
ollama run hf.co/Archsloth/Qwen3-4B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Archsloth/Qwen3-4B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Archsloth/Qwen3-4B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Archsloth/Qwen3-4B-GGUF with Docker Model Runner:
docker model run hf.co/Archsloth/Qwen3-4B-GGUF:Q4_K_M
- Lemonade
How to use Archsloth/Qwen3-4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Archsloth/Qwen3-4B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen3-4B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Archsloth/Qwen3-4B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Archsloth/Qwen3-4B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Archsloth/Qwen3-4B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Archsloth/Qwen3-4B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Archsloth/Qwen3-4B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
card: direct unsloth comparison + per-language and frontier charts
Browse files- README.md +46 -15
- chart_axes.svg +94 -0
- chart_frontier.svg +56 -0
README.md
CHANGED
|
@@ -42,6 +42,17 @@ tags:
|
|
| 42 |
<b>Most quantized weights ship with an adjective. Ours ship with a table.</b>
|
| 43 |
</p>
|
| 44 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
---
|
| 46 |
|
| 47 |
> ### 📚 Collection
|
|
@@ -80,23 +91,28 @@ There is no Q5 file here. [Why not.](#what-we-did-not-win)
|
|
| 80 |
|
| 81 |
## `[measured]` Q4_K_M — ten axes, 2,497,280,800 bytes on both sides
|
| 82 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
KL divergence from the bf16 original. **Lower is better.** `llama-perplexity --kl-divergence`,
|
| 84 |
ctx 512, 60 chunks for Korean and English, 40 for the rest. Every ARCHsloth number below comes
|
| 85 |
from **one single build**.
|
| 86 |
|
| 87 |
-
| | script | **ARCHsloth Q4** | `
|
| 88 |
-
|---|---|---|---|---|---|
|
| 89 |
-
| | file size | **2.4973 GB** | 2.4973 GB | 2.5463 GB | 2.4973 GB |
|
| 90 |
-
| Korean | Hangul | **0.024297** | 0.053290 | 0.046637 | 0.081506 |
|
| 91 |
-
| English | Latin | **0.032344** | 0.048443 | 0.043660 | 0.066203 |
|
| 92 |
-
| Spanish | Latin | **0.033679** | 0.048234 | 0.043093 | 0.069169 |
|
| 93 |
-
| Japanese | Kana·Han | **0.034148** | 0.051534 | 0.045666 | 0.075946 |
|
| 94 |
-
| Chinese | Han | **0.033088** | 0.048051 | 0.042510 | 0.073735 |
|
| 95 |
-
| Russian | Cyrillic | **0.032852** | 0.043831 | 0.038792 | 0.065831 |
|
| 96 |
-
| Arabic | Arabic | **0.040871** | 0.053212 | 0.047022 | 0.067085 |
|
| 97 |
-
| Hindi | Devanagari | **0.031617** | 0.039979 | 0.035678 | 0.053405 |
|
| 98 |
-
| Thai | Thai | **0.024561** | 0.031301 | 0.027782 | 0.045149 |
|
| 99 |
-
| **Source code** | — | **0.010191** | 0.021659 | 0.018216 | 0.039165 |
|
| 100 |
|
| 101 |
Ten axes, seven writing systems, ahead on every one — against the same-size file and against
|
| 102 |
the larger one. Lowest separation is 3.8 σ (Hindi); highest is 20.4 σ (Korean).
|
|
@@ -105,14 +121,29 @@ the larger one. Lowest separation is 3.8 σ (Hindi); highest is 20.4 σ (Korean)
|
|
| 105 |
not.** They are the control, so that a build which buys one language with another has somewhere
|
| 106 |
to show it.
|
| 107 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 108 |
### `[measured]` Q6_K and Q8_0
|
| 109 |
|
| 110 |
| | size | Korean | English |
|
| 111 |
|---|---|---|---|
|
| 112 |
| **ARCHsloth Q6_K** | 3.3063 GB | **0.004077** | **0.004022** |
|
| 113 |
-
| `
|
|
|
|
| 114 |
| **ARCHsloth Q8_0** | 4.2804 GB | **0.001624** | **0.001371** |
|
| 115 |
-
| `
|
|
|
|
| 116 |
|
| 117 |
Same-top-p and RMS Δp move the same direction on every row. Full statistics: [`EVAL.md`](EVAL.md).
|
| 118 |
Raw per-run logs: [`eval/logs/`](eval/logs).
|
|
|
|
| 42 |
<b>Most quantized weights ship with an adjective. Ours ship with a table.</b>
|
| 43 |
</p>
|
| 44 |
|
| 45 |
+
<p align="center">
|
| 46 |
+
<b>Head to head with <a href="https://huggingface.co/unsloth/Qwen3-4B-GGUF"><code>unsloth/Qwen3-4B-GGUF</code></a> at the identical
|
| 47 |
+
2,497,280,800 bytes — KL divergence from bf16, lower is better:</b><br>
|
| 48 |
+
Korean <b>−54.4 %</b> · Code <b>−52.9 %</b> ·
|
| 49 |
+
Japanese <b>−33.7 %</b> · English <b>−33.2 %</b> ·
|
| 50 |
+
Chinese <b>−31.1 %</b> · Spanish <b>−30.2 %</b> ·
|
| 51 |
+
Russian <b>−25.0 %</b> · Arabic <b>−23.2 %</b> ·
|
| 52 |
+
Thai <b>−21.5 %</b> · Hindi <b>−20.9 %</b><br>
|
| 53 |
+
<sub>Ten axes measured, ten ahead. Seven of them are not in our calibration set.</sub>
|
| 54 |
+
</p>
|
| 55 |
+
|
| 56 |
---
|
| 57 |
|
| 58 |
> ### 📚 Collection
|
|
|
|
| 91 |
|
| 92 |
## `[measured]` Q4_K_M — ten axes, 2,497,280,800 bytes on both sides
|
| 93 |
|
| 94 |
+
<p align="center">
|
| 95 |
+
<img src="chart_axes.svg" width="1000"
|
| 96 |
+
alt="KL divergence by language — ARCHsloth Q4 against unsloth Q4_K_M and UD-Q4_K_XL, ten axes">
|
| 97 |
+
</p>
|
| 98 |
+
|
| 99 |
KL divergence from the bf16 original. **Lower is better.** `llama-perplexity --kl-divergence`,
|
| 100 |
ctx 512, 60 chunks for Korean and English, 40 for the rest. Every ARCHsloth number below comes
|
| 101 |
from **one single build**.
|
| 102 |
|
| 103 |
+
| | script | **ARCHsloth Q4** | unsloth<br>`Q4_K_M` | unsloth<br>`UD-Q4_K_XL` | stock<br>`llama-quantize` | **vs unsloth**<br>same size |
|
| 104 |
+
|---|---|---|---|---|---|---|
|
| 105 |
+
| | file size | **2.4973 GB** | 2.4973 GB | 2.5463 GB | 2.4973 GB | — |
|
| 106 |
+
| Korean | Hangul | **0.024297** | 0.053290 | 0.046637 | 0.081506 | **−54.4 %** |
|
| 107 |
+
| English | Latin | **0.032344** | 0.048443 | 0.043660 | 0.066203 | **−33.2 %** |
|
| 108 |
+
| Spanish | Latin | **0.033679** | 0.048234 | 0.043093 | 0.069169 | **−30.2 %** |
|
| 109 |
+
| Japanese | Kana·Han | **0.034148** | 0.051534 | 0.045666 | 0.075946 | **−33.7 %** |
|
| 110 |
+
| Chinese | Han | **0.033088** | 0.048051 | 0.042510 | 0.073735 | **−31.1 %** |
|
| 111 |
+
| Russian | Cyrillic | **0.032852** | 0.043831 | 0.038792 | 0.065831 | **−25.0 %** |
|
| 112 |
+
| Arabic | Arabic | **0.040871** | 0.053212 | 0.047022 | 0.067085 | **−23.2 %** |
|
| 113 |
+
| Hindi | Devanagari | **0.031617** | 0.039979 | 0.035678 | 0.053405 | **−20.9 %** |
|
| 114 |
+
| Thai | Thai | **0.024561** | 0.031301 | 0.027782 | 0.045149 | **−21.5 %** |
|
| 115 |
+
| **Source code** | — | **0.010191** | 0.021659 | 0.018216 | 0.039165 | **−52.9 %** |
|
| 116 |
|
| 117 |
Ten axes, seven writing systems, ahead on every one — against the same-size file and against
|
| 118 |
the larger one. Lowest separation is 3.8 σ (Hindi); highest is 20.4 σ (Korean).
|
|
|
|
| 121 |
not.** They are the control, so that a build which buys one language with another has somewhere
|
| 122 |
to show it.
|
| 123 |
|
| 124 |
+
### `[measured]` The whole curve, not one point
|
| 125 |
+
|
| 126 |
+
<p align="center">
|
| 127 |
+
<img src="chart_frontier.svg" width="1000"
|
| 128 |
+
alt="File size against KL divergence — the ARCHsloth curve sits under the unsloth curve at every size">
|
| 129 |
+
</p>
|
| 130 |
+
|
| 131 |
+
Nine of those points are real files: six downloaded from
|
| 132 |
+
[`unsloth/Qwen3-4B-GGUF`](https://huggingface.co/unsloth/Qwen3-4B-GGUF), three built here. Nothing is interpolated.
|
| 133 |
+
Their curve includes the `UD-*-XL` files, which are **larger** than the same-named standard
|
| 134 |
+
file — and our curve is still below theirs. Read it vertically for "same size, closer to the
|
| 135 |
+
original"; read it horizontally for "same quality, smaller file".
|
| 136 |
+
|
| 137 |
### `[measured]` Q6_K and Q8_0
|
| 138 |
|
| 139 |
| | size | Korean | English |
|
| 140 |
|---|---|---|---|
|
| 141 |
| **ARCHsloth Q6_K** | 3.3063 GB | **0.004077** | **0.004022** |
|
| 142 |
+
| unsloth `Q6_K` | 3.3063 GB | 0.006969 | 0.005765 |
|
| 143 |
+
| | | **−41.5 %** | **−30.2 %** |
|
| 144 |
| **ARCHsloth Q8_0** | 4.2804 GB | **0.001624** | **0.001371** |
|
| 145 |
+
| unsloth `Q8_0` | 4.2804 GB | 0.001827 | 0.001502 |
|
| 146 |
+
| | | **−11.1 %** | **−8.7 %** |
|
| 147 |
|
| 148 |
Same-top-p and RMS Δp move the same direction on every row. Full statistics: [`EVAL.md`](EVAL.md).
|
| 149 |
Raw per-run logs: [`eval/logs/`](eval/logs).
|
chart_axes.svg
ADDED
|
|
chart_frontier.svg
ADDED
|
|