Is it possible to offload n-gram to an NVMe SSD?

#11
by lingyezhixing - opened

Thanks to 0Day's support, I noticed that unsloth has taken over the adaptation of Qwen4 for llama.cpp. It seems that n-gram queries are O(1). Could it be possible to offload them to memory or even to disk to reduce GPU memory usage?

From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.

Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.

Are the ngram layers necessary, does it just run slower without it (like without ngram speculative decoding), or will it simply won't work?

From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.

Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.

Some one told me FP8 engram is 51GB so if some one has a 32GB ram they are cooked? What about swap file?

Swap is on the ssd, so falls into the same pitfall as ssd streaming. You can maybe have it working if ngrams are fp4/int4, but how well ngrams quantize, afaik, is still under-explored.

Are the ngram layers necessary, does it just run slower without it (like without ngram speculative decoding), or will it simply won't work?

I saw someone perform some testing on the subject of 'just skipping' the ngrams: agentionai/Qwen3.8-Flash-Next-ROCmFP4-FAST-GGUF . There are some reports in the model card there. The tldr is that, it hurts the model a lot.

From what I understand of how the architecture works, no. There is some discussion about this in llama.cpp PRs, but it boils down to the fact that the n-gram is hashed, and therefore its unpredictable what/where needs to be read. Hence why random access memory (RAM) is perfect for it, because unlike an ssd, which (ideally) wants stuff read in a 'specific order', RAM doesn't care for the order. The hashing part of the n-gram makes it essentially "you have to stream ~1GB per token" if you try to do it from an ssd, and even at 7gb/s speeds (i think that is roughly what current top of the line is?) that means optimally about 7 t/s.

Could be wrong, I'm trying to piece it together myself, but it seems unlikely from my initial research. I do hope I'm wrong.

After some research, I think traditional mmap might not work—we'd have to bypass the system entirely and directly read the n-gram tables. The main network would run on the GPU, while the n-gram operations would be offloaded to NVMe. In this setup, we could leverage NVMe's 4K random read performance to achieve O(1) lookup times.

During the prefill phase, since all tokens are known upfront, we can first deduplicate them, then asynchronously and in parallel fully utilize NVMe's performance while simultaneously performing lookup operations and main network computations. In the decoding phase, after each token is decoded, we immediately perform asynchronous and parallel n-gram lookups and computations on the first two layers, which then merge at the third layer—this is roughly the logic. It could still be optimized further by caching frequently used n-grams in memory for faster access.

Honestly, I don't think anyone would actually implement this. All of this would require bypassing the page cache, otherwise it would be bottlenecked by the filesystem. Llama currently lacks direct file reading logic and asynchronous inference capabilities. Adding two complex features—direct file access and asynchronous inference—to a single model architecture goes against Llama's original design philosophy. I hope there's a better solution out there.

I need to investigate more, but I think that ngram block is a cache.
So removing rarely used parts to make it smaller to fit into RAM might work.
Removing all of it lobotomizes the model.

I don't feel too optimistic about removing parts of it, as to my understanding it's a lot more "specific knowledge is in a specific block" than experts, and that's why, I suspect, it won't be as forgiving as REAP-ing experts (where knowledge is more spread/generalized over many experts). But if you have a little free time, I'd be curious to see your findings on the topic! Many people with 32gb ram would probably benefit a lot from a 50% trim, I'm just not too optimistic on the effect. But adapting REAP for the ngram sounds fun.

Encoding the ngram to IQ3_S should work.
It has the block scales and decent resolution.
Needs code updates but 22GB is possible here!

With some work it'll fit on a single 5090 and 32GB RAM

Just saw that unsloth already quantized to below 32GB: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/tree/main/UD-IQ1_S
26.82 GiB

I see 72.5 GB there...

We were talking about the ngram table that will sit in RAM. The regular tensors will go to VRAM

Got SSD streaming working on Mac with Metal and Metal I/O.

No performance loss on inference from what I can tell with this approach. 36t/s decode at 0 context size for both resident and SSD offloaded n-gram. Prefill is slower with offload.

https://github.com/ggml-org/llama.cpp/commit/b591584708a6d4fff43c08cb89809e88b389cb08


Screenshot 2026-08-26 at 23.54.38
70gb memory usage on 128gb macbook, 25t/s at 55k context

I am very happy to be wrong, it does seem like that commit is indeed working. I had to port it to cuda, so that it ran on my DGX Spark. I am not going to commit anything, as I just threw glm-5.3-flash at the task, but as a PoC it seems to actually work fine. SSD offloading seems like may end up being the way to run it. I am noticing a ~10-15% hit to prefil and decode, the latter should be fine once MTP works, prefill is still 560 t/s on the spark, so not terrible.

For Strix Halo users, you can get the N-gram table on disk with strix-llama.cpp: https://github.com/halo-box/strix-llama.cpp#what-differs-from-upstream
With it, I run UD-Q4_K_XL at around 40 t/s on my 128 GB Strix Halo.

Sign up or log in to comment