Instructions to use unsloth/Qwen3.8-Flash-Next-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/Qwen3.8-Flash-Next-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use unsloth/Qwen3.8-Flash-Next-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "unsloth/Qwen3.8-Flash-Next-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/Qwen3.8-Flash-Next-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
- Ollama
How to use unsloth/Qwen3.8-Flash-Next-GGUF with Ollama:
ollama run hf.co/unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use unsloth/Qwen3.8-Flash-Next-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/Qwen3.8-Flash-Next-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/Qwen3.8-Flash-Next-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use unsloth/Qwen3.8-Flash-Next-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/Qwen3.8-Flash-Next-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Is it possible to offload n-gram to an NVMe SSD?
Thanks to 0Day's support, I noticed that unsloth has taken over the adaptation of Qwen4 for llama.cpp. It seems that n-gram queries are O(1). Could it be possible to offload them to memory or even to disk to reduce GPU memory usage?
From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.
Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.
Are the ngram layers necessary, does it just run slower without it (like without ngram speculative decoding), or will it simply won't work?
From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.
Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.
Some one told me FP8 engram is 51GB so if some one has a 32GB ram they are cooked? What about swap file?
Swap is on the ssd, so falls into the same pitfall as ssd streaming. You can maybe have it working if ngrams are fp4/int4, but how well ngrams quantize, afaik, is still under-explored.
Are the ngram layers necessary, does it just run slower without it (like without ngram speculative decoding), or will it simply won't work?
I saw someone perform some testing on the subject of 'just skipping' the ngrams: agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF . There are some reports in the model card there. The tldr is that, it hurts the model a lot.
From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.
Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.
After some research, I think traditional mmap might not work—we'd have to bypass the system entirely and directly read the n-gram tables. The main network would run on the GPU, while the n-gram operations would be offloaded to NVMe. In this setup, we could leverage NVMe's 4K random read performance to achieve O(1) lookup times.
During the prefill phase, since all tokens are known upfront, we can first deduplicate them, then asynchronously and in parallel fully utilize NVMe's performance while simultaneously performing lookup operations and main network computations. In the decoding phase, after each token is decoded, we immediately perform asynchronous and parallel n-gram lookups and computations on the first two layers, which then merge at the third layer—this is roughly the logic. It could still be optimized further by caching frequently used n-grams in memory for faster access.
Honestly, I don't think anyone would actually implement this. All of this would require bypassing the page cache, otherwise it would be bottlenecked by the filesystem. Llama currently lacks direct file reading logic and asynchronous inference capabilities. Adding two complex features—direct file access and asynchronous inference—to a single model architecture goes against Llama's original design philosophy. I hope there's a better solution out there.
I need to investigate more, but I think that ngram block is a cache.
So removing rarely used parts to make it smaller to fit into RAM might work.
Removing all of it lobotomizes the model.
I don't feel too optimistic about removing parts of it, as to my understanding it's a lot more "specific knowledge is in a specific block" than experts, and that's why, I suspect, it won't be as forgiving as REAP-ing experts (where knowledge is more spread/generalized over many experts). But if you have a little free time, I'd be curious to see your findings on the topic! Many people with 32gb ram would probably benefit a lot from a 50% trim, I'm just not too optimistic on the effect. But adapting REAP for the ngram sounds fun.
Encoding the ngram to IQ3_S should work.
It has the block scales and decent resolution.
Needs code updates but 22GB is possible here!
With some work it'll fit on a single 5090 and 32GB RAM
Just saw that unsloth already quantized to below 32GB: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main/UD-IQ1_S
26.82 GiB
I see 72.5 GB there...
We were talking about the ngram table that will sit in RAM. The regular tensors will go to VRAM
Got SSD streaming working on Mac with Metal and Metal I/O.
No performance loss on inference from what I can tell with this approach. 36t/s decode at 0 context size for both resident and SSD offloaded n-gram. Prefill is slower with offload.
https://github.com/ggml-org/llama.cpp/commit/b591584708a6d4fff43c08cb89809e88b389cb08
I am very happy to be wrong, it does seem like that commit is indeed working. I had to port it to cuda, so that it ran on my DGX Spark. I am not going to commit anything, as I just threw glm-5.3-flash at the task, but as a PoC it seems to actually work fine. SSD offloading seems like may end up being the way to run it. I am noticing a ~10-15% hit to prefil and decode, the latter should be fine once MTP works, prefill is still 560 t/s on the spark, so not terrible.
For Strix Halo users, you can get the N-gram table on disk with strix-llama.cpp: https://github.com/halo-box/strix-llama.cpp#what-differs-from-upstream
With it, I run UD-Q4_K_XL at around 40 t/s on my 128 GB Strix Halo.
