Instructions to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- VibeVoice
How to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with VibeVoice:
import torch, soundfile as sf, librosa, numpy as np from vibevoice.processor.vibevoice_processor import VibeVoiceProcessor from vibevoice.modular.modeling_vibevoice_inference import VibeVoiceForConditionalGenerationInference # Load voice sample (should be 24kHz mono) voice, sr = sf.read("path/to/voice_sample.wav") if voice.ndim > 1: voice = voice.mean(axis=1) if sr != 24000: voice = librosa.resample(voice, sr, 24000) processor = VibeVoiceProcessor.from_pretrained("EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF") model = VibeVoiceForConditionalGenerationInference.from_pretrained( "EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF", torch_dtype=torch.bfloat16 ).to("cuda").eval() model.set_ddpm_inference_steps(5) inputs = processor(text=["Speaker 0: Hello!\nSpeaker 1: Hi there!"], voice_samples=[[voice]], return_tensors="pt") audio = model.generate(**inputs, cfg_scale=1.3, tokenizer=processor.tokenizer).speech_outputs[0] sf.write("output.wav", audio.cpu().numpy().squeeze(), 24000) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
Use Docker
docker model run hf.co/EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
- LM Studio
- Jan
- Ollama
How to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with Ollama:
ollama run hf.co/EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
- Unsloth Desktop
- Docker Model Runner
How to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with Docker Model Runner:
docker model run hf.co/EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
- Lemonade
How to use EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF:Q8_0
Run and chat with the model
lemonade run user.VibeVoice-ASR-Streaming-7B-GGUF-Q8_0
List all available models
lemonade list
- Atomic Chat
VibeVoice-ASR-Streaming-7B GGUF
Unofficial GGUF conversion of
microsoft/VibeVoice-ASR-Streaming-7B
for the native Rust implementation in
Liyulingyue/rust-model-inference.
The model performs streaming, speaker-attributed automatic speech recognition
for Chinese, English, French, German, Italian, Japanese, Korean, Portuguese,
Russian, and Spanish. It can also use context terms supplied through the Rust
CLI's --prompt option.
Files
Both files are required for ASR inference.
| File | Size | Contents |
|---|---|---|
VibeVoice-ASR-Streaming-7B-Q8_0.gguf |
7.54 GiB | Qwen2.5 decoder: 198 Q8_0 tensors and 141 F32 tensors |
mmproj-VibeVoice-ASR-Streaming-7B-BF16.gguf |
1.33 GiB | Acoustic/semantic encoders and speech connectors: 562 BF16 tensors |
The LLM has 28 layers, a hidden size of 3584, a context length of 131072, and
a vocabulary size of 152064. The mmproj expects 24 kHz audio and uses a custom
vibevoice_asr audio projector.
Compatibility
The validated runtime is rust-model-inference after merged
PR #47.
The LLM uses llama.cpp-style Qwen2 tensor names, but the end-to-end ASR model is not compatible with stock llama.cpp. The audio encoder/projector is a VibeVoice-specific mmproj that requires the Rust implementation above.
Rust Usage
git clone https://github.com/Liyulingyue/rust-model-inference.git
cd rust-model-inference
cargo run --release --bin rust-model-inference -- \
--model /path/to/VibeVoice-ASR-Streaming-7B-Q8_0.gguf \
--mmproj /path/to/mmproj-VibeVoice-ASR-Streaming-7B-BF16.gguf \
--audio /path/to/input.wav \
--temp 0 \
--max-tokens 128 \
--threads 8 \
--out transcription.txt
The current CLI accepts PCM16 WAV input. It downmixes multi-channel input and
resamples it to the model's 24 kHz sample rate. --max-tokens is the maximum
number of generated tokens per streaming chunk.
Use --prompt to provide names or domain terms as recognition context:
--prompt "Microsoft,VibeVoice"
For reproducible encoder latents, set:
export VIBEVOICE_ASR_DETERMINISTIC=1
Conversion
The converter reads the original sharded BF16 safetensors through mmap and does not require PyTorch:
python3 tools/vibevoice/convert_vibevoice_asr.py \
/path/to/VibeVoice-ASR-Streaming-7B \
--out-dir /path/to/output
The language decoder is quantized to GGML Q8_0. The acoustic and semantic encoders and connectors remain BF16. The unused acoustic tokenizer decoder is not included because ASR does not synthesize audio.
Alignment
The Rust implementation was checked against the original BF16 safetensors on a fixed prompt and audio fixture:
| Checkpoint | Result |
|---|---|
| Decoder layer 0 hidden NRMSE | 0.005884 |
| Worst decoder layer hidden NRMSE (layer 25) | 0.023366 |
| Decoder layer 27 hidden NRMSE | 0.015076 |
| Final normalized hidden NRMSE | 0.020717 |
| Full 152064-value logits NRMSE | 0.009781 |
| Maximum absolute logits error | 0.196437 |
The top-8 token ordering matched exactly:
[58, 151665, 715, 42474, 39379, 32622, 43504, 37073]
The BF16 audio encoder Oracle covered 12 checkpoints with a maximum absolute error of 0.000070.
These are numerical comparisons, not bitwise equality. Q8_0 weight and activation quantization necessarily differs from the original BF16 model. The decoder comparison covers the final sequence row at every layer for one fixed fixture; it does not claim exhaustive equality for every input or internal Q/K/V tensor.
Checksums
255ca05bb3b4f34ab02922da5bc3f4ac6de489552175cc013c1a1db7e69d9514 VibeVoice-ASR-Streaming-7B-Q8_0.gguf
020e256a28f9ccb1e0ad658467545a6a81bd9d3ca2f52cb4ba146712afc1699b mmproj-VibeVoice-ASR-Streaming-7B-BF16.gguf
Limitations
- Only the CPU inference path has been validated for this conversion.
- Quantization can change recognition output relative to the BF16 checkpoint.
- The Rust CLI currently requires PCM16 WAV files; microphone/WebSocket serving is not included in this model repository.
- Review the original model card for intended use, evaluation, and responsible use guidance.
License and Attribution
The source model is provided by Microsoft under the MIT License. This
repository redistributes converted model weights under the same license. See
LICENSE and the
original model card.
- Downloads last month
- 372
8-bit
Model tree for EvoAwaken-Workshop/VibeVoice-ASR-Streaming-7B-GGUF
Base model
microsoft/VibeVoice-ASR-Streaming-7B