Team Ai
Modelpublic

mudler/vibevoice.cpp-models

sourceHugging Facemitupdated 6mo agoView on Hugging Face
12likes14kdownloads
README.md85 linesDownload Raw Back to root
1---2license: mit3library_name: vibevoice.cpp4tags:5  - tts6  - asr7  - speech8  - vibevoice9  - gguf10  - ggml11base_model:12  - microsoft/VibeVoice-Realtime-0.5B13  - microsoft/VibeVoice-ASR14---15 16# vibevoice.cpp — quantized model bundle17 18**Brought to you by the [LocalAI](https://github.com/mudler/LocalAI) team** — the creators of LocalAI, the open-source AI engine that runs any model — LLMs, vision, voice, image, video — on any hardware. No GPU required.19 20Quantized GGUF weights for [vibevoice.cpp](https://github.com/mudler/vibevoice.cpp),21a C/C++ port of Microsoft VibeVoice (TTS + ASR) on top of `ggml`.22 23| File | Source | Quant | Size |24| ---- | ------ | ----- | ---- |25| `vibevoice-realtime-0.5B-q8_0.gguf` | `microsoft/VibeVoice-Realtime-0.5B` | Q8_0 (matmul) + F16 | ~1.6 GB |26| `vibevoice-asr-q8_0.gguf`           | `microsoft/VibeVoice-ASR`           | Q8_0 (matmul) + F16 | ~13 GB |27| `voice-en-Carter_man.gguf`          | upstream voice prompt cache         | F16                  | 8 MB |28| `voice-en-Emma.gguf`                | upstream voice prompt cache         | F16                  | 6 MB |29| `tokenizer.gguf`                    | Qwen2.5 BPE + VibeVoice specials    | —                    | 6 MB |30 31## Quantization scheme32 33`scripts/quantize_gguf.py` in the source repo selectively quantizes only the34LM matmul weights — attention q/k/v/o, ffn gate/up/down, and lm_head — to35Q8_0. Everything else (1-D conv kernels, RMSNorm scales, biases,36layer-scale gammas, token embeddings, small scalars) passes through37unchanged. The conv1d implementation in vibevoice.cpp casts kernels to F1638inline rather than dequantizing on the fly, so quantizing those would39corrupt the convolution outputs.40 41Q8_0 was chosen because it's pure-Python implementable in `gguf-py` and42gives a ~60% size reduction on the 7B ASR model with no measurable43quality regression in the closed-loop TTS → ASR roundtrip test.44 45## Quickstart46 47```bash48git clone --recursive https://github.com/mudler/vibevoice.cpp49cd vibevoice.cpp && cmake -B build -DVIBEVOICE_BUILD_TESTS=ON && cmake --build build -j50 51# Pull this bundle52mkdir -p models && cd models53hf download mudler/vibevoice.cpp-models --local-dir .54cd ..55 56# TTS57build/bin/vibevoice-cli tts \58    --model models/vibevoice-realtime-0.5B-q8_0.gguf \59    --voice models/voice-en-Carter_man.gguf \60    --tokenizer models/tokenizer.gguf \61    --text "Hello world this is a test of the synthesis system." \62    --out hello.wav63 64# ASR65build/bin/vibevoice-cli asr \66    --model models/vibevoice-asr-q8_0.gguf \67    --tokenizer models/tokenizer.gguf \68    --audio hello.wav69# -> [{"Start":0,"End":2.8,"Speaker":0,"Content":"Hello world, this is a test of the synthesis system."}]70```71 72## Closed-loop verification73 74The `test_closed_loop` ctest in vibevoice.cpp runs TTS → ASR end-to-end75and asserts ≥80% source-word recall in the recovered transcript. With76this bundle (both Q8_0 models) it passes at 10/10 (100 %).77 78## License79 80Weights are derived from Microsoft VibeVoice81([VibeVoice-Realtime-0.5B](https://huggingface.co/microsoft/VibeVoice-Realtime-0.5B)82and [VibeVoice-ASR](https://huggingface.co/microsoft/VibeVoice-ASR));83follow the upstream model licenses for use. The conversion + quantization84tooling is released under MIT as part of vibevoice.cpp.85