NVIDIA NemotronLabs VoiceChat 11B - GGUF
Custom GGUF quantizations of NVIDIA/NemotronLabs-VoiceChat-11B.
These files were converted from the original NVIDIA SafeTensors checkpoint to GGUF and then quantized using custom GGML-based C++ workers.
Running these files
llama.cpp now runs this model, through the fork
sansamour/llama-voicechat.cpp.
Speech in, speech out, one duplex timeline, tool calls, on CPU or CUDA. No torch,
no NeMo.
The single-container GGUFs below cannot be loaded as they are: they carry NeMo
tensor names, no llama.cpp KV metadata and no tokenizer. Three converters in the
fork split one of them into the four files the runtime wants, and the result of
doing that to nemotron_voicechat_11b-Q4_0.gguf is in the llamacpp/ folder
of this repo:
| file | size | what |
|---|---|---|
llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0.gguf |
4.67 GiB | the language model, a stock nemotron_h |
llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0-function-head.gguf |
315 MiB | the turn-taking / tool-call head, which llama.cpp would reject inside the model file |
llamacpp/mmproj-voicechat-perception-Q4_0.gguf |
435 MiB | the causal FastConformer speech encoder, as an mtmd audio projector |
llamacpp/voicechat-tts-Q4_0.gguf |
686 MiB | the speech generator: gemma3 backbone, MoG head, 31 stage RVQ, codec decoder |
hf download hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF --include "llamacpp/*" --local-dir .
llama-voicechat -m llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0.gguf --mmproj llamacpp/mmproj-voicechat-perception-Q4_0.gguf --tts llamacpp/voicechat-tts-Q4_0.gguf --audio question.wav --tts-out answer.wav
Keep the four files in one folder; the function head is found by name next to the
model. Set VC_NO_BARGE=1 and VC_FORCE_BOS=1 for anything push-to-talk - the
model is duplex, so left alone it opens its own turn about a second into the clip
and answers only that first second.
Nothing was requantized on the way. A Q4_0 tensor in the source is copied out block for block and stays bit-identical.
The model answers in English only.
Original model
- Base model:
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B - Model: NVIDIA NemotronLabs VoiceChat 11B
- Format of original checkpoint: SafeTensors
- Original tensor count: 1632
- GGUF version: GGUF V3
- Conversion target: GGUF
The original model and its configuration are available at:
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
Quantized files
| File | Quantization | Size |
|---|---|---|
nemotron_voicechat_11b-Q8_0.gguf |
Q8_0 | ~12.04 GB |
nemotron_voicechat_11b-Q4_0.gguf |
Q4_0 | ~6.62 GB |
Actual files produced during conversion:
Q8_0
Output size: 12,035,312,704 bytes
Final size : 11.2088 GiB
Q8_0 tensors : 1369
Kept tensors : 263
F16 kept : 257
Non-F16 kept : 6
Q4_0
Output size: 6,618,150,208 bytes
Final size : 6.16363 GiB
Q4_0 tensors : 1369
F16 kept : 257
Non-F16 kept : 6
Conversion pipeline
The conversion was performed entirely locally.
NVIDIA SafeTensors
|
v
model.safetensors
|
| custom Python converter
v
nemotron_voicechat_11b-f16.gguf
|
+----------------------------+
| |
| custom C++ GGML worker | custom C++ GGML worker
v v
Q8_0 Q4_0
| |
v v
nemotron_voicechat_11b- nemotron_voicechat_11b-
Q8_0.gguf Q4_0.gguf
The F16 intermediate GGUF contained:
1632 tensors
GGUF V3
The quantization workers preserve:
- tensor order
- tensor names
- tensor dimensions
- GGUF metadata
- custom VoiceChat metadata
- non-F16 tensors
Only eligible F16 tensors are quantized.
Q8_0 conversion
The Q8_0 conversion uses a custom C++ worker based on GGML quantization routines.
The policy used for this build is:
F16 tensor
|
+-- compatible with Q8_0
| |
| +--> Q8_0
|
+-- incompatible
| |
| +--> keep F16
|
+-- non-F16
|
+--> keep original type
Result:
1369 tensors -> Q8_0
257 tensors -> F16
6 tensors -> original non-F16 type
Q4_0 conversion
The Q4_0 conversion uses a separate custom C++ worker.
Policy:
F16 tensor
|
+-- compatible with Q4_0
| |
| +--> Q4_0
|
+-- incompatible
| |
| +--> keep F16
|
+-- non-F16
|
+--> keep original type
Result:
1369 tensors -> Q4_0
257 tensors -> F16
6 tensors -> original non-F16 type
Compatibility
An unmodified llama.cpp binary cannot run these files. llama-quantize cannot
touch them either:
unknown model architecture: 'nemotron_voicechat'
That is expected: this is one container holding five models under NeMo tensor
names, and llama.cpp has no such architecture. The fork
sansamour/llama-voicechat.cpp
does not add one - it splits the container into pieces llama.cpp already
understands (a nemotron_h model, an mtmd audio projector, and two side GGUFs
read by the tool itself). See Running these files above;
the converted set is in llamacpp/.
Hardware targets
Approximate target sizes:
Q8_0
~12 GB model file.
Suitable for testing on GPUs with approximately:
- 16 GB VRAM
- 12 GB VRAM with CPU/RAM offload
- larger GPUs
Q4_0
~6.6 GB model file.
Suitable for testing on GPUs with approximately:
- 8 GB VRAM
- 12 GB VRAM
- 16 GB VRAM
- larger GPUs
Actual runtime VRAM usage will be higher than the GGUF file size because of:
- KV cache
- activations
- temporary buffers
- runtime workspace
- audio processing buffers
- CUDA kernels / backend allocations
Therefore the GGUF file size should NOT be interpreted as the exact VRAM requirement.
Quantization quality
These are custom quantizations.
No claim is made here that Q8_0 or Q4_0 produces identical output quality to the original NVIDIA SafeTensors model.
Users should benchmark:
- speech recognition accuracy
- audio understanding
- voice response quality
- latency
- memory usage
against the original model.
Original NVIDIA model
Please refer to the original NVIDIA repository for:
- model architecture
- configuration
- intended usage
- license
- model documentation
- original model limitations
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
License
The license and usage conditions of the original NVIDIA model apply.
Please consult the original model repository before using or redistributing these files:
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
Disclaimer
This repository is an unofficial community conversion.
It is not an official NVIDIA release.
The quantized GGUF files were produced independently from the original NVIDIA checkpoint and are provided for research, experimentation and local inference development.
Credits
Original model:
NVIDIA NemotronLabs VoiceChat 11B
Original repository:
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
GGUF conversion and quantization:
hoidhxd
Repository:
https://huggingface.co/hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF
- Downloads last month
- 2,113
4-bit
8-bit
16-bit
Model tree for hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF
Base model
nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base