NVIDIA NemotronLabs VoiceChat 11B - GGUF

Custom GGUF quantizations of NVIDIA/NemotronLabs-VoiceChat-11B.

These files were converted from the original NVIDIA SafeTensors checkpoint to GGUF and then quantized using custom GGML-based C++ workers.

Running these files

llama.cpp now runs this model, through the fork sansamour/llama-voicechat.cpp. Speech in, speech out, one duplex timeline, tool calls, on CPU or CUDA. No torch, no NeMo.

The single-container GGUFs below cannot be loaded as they are: they carry NeMo tensor names, no llama.cpp KV metadata and no tokenizer. Three converters in the fork split one of them into the four files the runtime wants, and the result of doing that to nemotron_voicechat_11b-Q4_0.gguf is in the llamacpp/ folder of this repo:

file size what
llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0.gguf 4.67 GiB the language model, a stock nemotron_h
llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0-function-head.gguf 315 MiB the turn-taking / tool-call head, which llama.cpp would reject inside the model file
llamacpp/mmproj-voicechat-perception-Q4_0.gguf 435 MiB the causal FastConformer speech encoder, as an mtmd audio projector
llamacpp/voicechat-tts-Q4_0.gguf 686 MiB the speech generator: gemma3 backbone, MoG head, 31 stage RVQ, codec decoder
hf download hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF --include "llamacpp/*" --local-dir .
llama-voicechat -m llamacpp/nemotron_voicechat_11b-stt-llm-Q4_0.gguf --mmproj llamacpp/mmproj-voicechat-perception-Q4_0.gguf --tts llamacpp/voicechat-tts-Q4_0.gguf --audio question.wav --tts-out answer.wav

Keep the four files in one folder; the function head is found by name next to the model. Set VC_NO_BARGE=1 and VC_FORCE_BOS=1 for anything push-to-talk - the model is duplex, so left alone it opens its own turn about a second into the clip and answers only that first second.

Nothing was requantized on the way. A Q4_0 tensor in the source is copied out block for block and stays bit-identical.

The model answers in English only.

Original model

  • Base model: nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
  • Model: NVIDIA NemotronLabs VoiceChat 11B
  • Format of original checkpoint: SafeTensors
  • Original tensor count: 1632
  • GGUF version: GGUF V3
  • Conversion target: GGUF

The original model and its configuration are available at:

https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

Quantized files

File Quantization Size
nemotron_voicechat_11b-Q8_0.gguf Q8_0 ~12.04 GB
nemotron_voicechat_11b-Q4_0.gguf Q4_0 ~6.62 GB

Actual files produced during conversion:

Q8_0

Output size: 12,035,312,704 bytes
Final size : 11.2088 GiB

Q8_0 tensors : 1369
Kept tensors : 263
F16 kept     : 257
Non-F16 kept : 6

Q4_0

Output size: 6,618,150,208 bytes
Final size : 6.16363 GiB

Q4_0 tensors : 1369
F16 kept     : 257
Non-F16 kept : 6

Conversion pipeline

The conversion was performed entirely locally.

NVIDIA SafeTensors
        |
        v
model.safetensors
        |
        | custom Python converter
        v
nemotron_voicechat_11b-f16.gguf
        |
        +----------------------------+
        |                            |
        | custom C++ GGML worker     | custom C++ GGML worker
        v                            v
Q8_0                           Q4_0
        |                            |
        v                            v
nemotron_voicechat_11b-       nemotron_voicechat_11b-
Q8_0.gguf                      Q4_0.gguf

The F16 intermediate GGUF contained:

1632 tensors
GGUF V3

The quantization workers preserve:

  • tensor order
  • tensor names
  • tensor dimensions
  • GGUF metadata
  • custom VoiceChat metadata
  • non-F16 tensors

Only eligible F16 tensors are quantized.

Q8_0 conversion

The Q8_0 conversion uses a custom C++ worker based on GGML quantization routines.

The policy used for this build is:

F16 tensor
    |
    +-- compatible with Q8_0
    |       |
    |       +--> Q8_0
    |
    +-- incompatible
    |       |
    |       +--> keep F16
    |
    +-- non-F16
            |
            +--> keep original type

Result:

1369 tensors -> Q8_0
257 tensors   -> F16
6 tensors     -> original non-F16 type

Q4_0 conversion

The Q4_0 conversion uses a separate custom C++ worker.

Policy:

F16 tensor
    |
    +-- compatible with Q4_0
    |       |
    |       +--> Q4_0
    |
    +-- incompatible
    |       |
    |       +--> keep F16
    |
    +-- non-F16
            |
            +--> keep original type

Result:

1369 tensors -> Q4_0
257 tensors   -> F16
6 tensors     -> original non-F16 type

Compatibility

An unmodified llama.cpp binary cannot run these files. llama-quantize cannot touch them either:

unknown model architecture: 'nemotron_voicechat'

That is expected: this is one container holding five models under NeMo tensor names, and llama.cpp has no such architecture. The fork sansamour/llama-voicechat.cpp does not add one - it splits the container into pieces llama.cpp already understands (a nemotron_h model, an mtmd audio projector, and two side GGUFs read by the tool itself). See Running these files above; the converted set is in llamacpp/.

Hardware targets

Approximate target sizes:

Q8_0

~12 GB model file.

Suitable for testing on GPUs with approximately:

  • 16 GB VRAM
  • 12 GB VRAM with CPU/RAM offload
  • larger GPUs

Q4_0

~6.6 GB model file.

Suitable for testing on GPUs with approximately:

  • 8 GB VRAM
  • 12 GB VRAM
  • 16 GB VRAM
  • larger GPUs

Actual runtime VRAM usage will be higher than the GGUF file size because of:

  • KV cache
  • activations
  • temporary buffers
  • runtime workspace
  • audio processing buffers
  • CUDA kernels / backend allocations

Therefore the GGUF file size should NOT be interpreted as the exact VRAM requirement.

Quantization quality

These are custom quantizations.

No claim is made here that Q8_0 or Q4_0 produces identical output quality to the original NVIDIA SafeTensors model.

Users should benchmark:

  • speech recognition accuracy
  • audio understanding
  • voice response quality
  • latency
  • memory usage

against the original model.

Original NVIDIA model

Please refer to the original NVIDIA repository for:

  • model architecture
  • configuration
  • intended usage
  • license
  • model documentation
  • original model limitations

https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

License

The license and usage conditions of the original NVIDIA model apply.

Please consult the original model repository before using or redistributing these files:

https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

Disclaimer

This repository is an unofficial community conversion.

It is not an official NVIDIA release.

The quantized GGUF files were produced independently from the original NVIDIA checkpoint and are provided for research, experimentation and local inference development.

Credits

Original model:

NVIDIA NemotronLabs VoiceChat 11B

Original repository:

https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

GGUF conversion and quantization:

hoidhxd

Repository:

https://huggingface.co/hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF

Downloads last month
2,113
GGUF
Model size
11B params
Architecture
nemotron_voicechat
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF

Space using hoidhxd/NVIDIA-NemotronLabs-VoiceChat-11B-GGUF 1