Team Ai
Modelpublic

mudler/parakeet-cpp-gguf

sourceHugging Faceotherupdated 2d agoView on Hugging Face
71likes25kdownloads
Model Card

Parakeet GGUF — models for parakeet.cpp

GGUF-format weights for parakeet.cpp, a C++/ggml port of NVIDIA NeMo Parakeet that matches the upstream PyTorch models on CPU. This single repo collects every supported model × quantization as a flat set of .gguf files — download just the one you need.

F16 is the recommended default — same accuracy as F32, ~1.7× smaller, and typically the fastest on modern CPUs via ggml's F32×F16 matmul fast path.

Models

tdt_ctc-110m

Source: nvidia/parakeet-tdt_ctc-110m · Hybrid TDT+CTC (FastConformer) · heads: TDT + CTC

FileVariantSizeWER vs NeMo
tdt_ctc-110m-f16.gguf ← recommendedF16267.5 MB0.0000
tdt_ctc-110m-q8_0.ggufQ8_0177.8 MB0.0000
tdt_ctc-110m-q6_k.ggufQ6_K155.9 MBnot measured
tdt_ctc-110m-q5_k.ggufQ5_K143.3 MBnot measured
tdt_ctc-110m-q4_k.ggufQ4_K131.4 MB0.0000

realtimeeou120m-v1

Source: nvidia/parakeet_realtime_eou_120m-v1 · Cache-aware streaming RNNT (FastConformer, EOU/EOB) · heads: RNNT (streaming) · License: NVIDIA Open Model License

FileVariantSizeWER vs NeMo
realtime_eou_120m-v1-f16.gguf ← recommendedF16266.5 MBnot measured
realtime_eou_120m-v1-q8_0.ggufQ8_0176.0 MBnot measured
realtime_eou_120m-v1-q6_k.ggufQ6_K153.9 MBnot measured
realtime_eou_120m-v1-q5_k.ggufQ5_K141.2 MBnot measured
realtime_eou_120m-v1-q4_k.ggufQ4_K129.1 MBnot measured

ctc-0.6b

Source: nvidia/parakeet-ctc-0.6b · CTC (FastConformer) · heads: CTC

FileVariantSizeWER vs NeMo
ctc-0.6b-f16.gguf ← recommendedF161373.4 MB0.0000
ctc-0.6b-q8_0.ggufQ8_0875.4 MB0.0000
ctc-0.6b-q6_k.ggufQ6_K746.8 MBnot measured
ctc-0.6b-q5_k.ggufQ5_K676.3 MBnot measured
ctc-0.6b-q4_k.ggufQ4_K609.9 MBnot measured

rnnt-0.6b

Source: nvidia/parakeet-rnnt-0.6b · RNNT transducer (FastConformer) · heads: RNNT

FileVariantSizeWER vs NeMo
rnnt-0.6b-f16.gguf ← recommendedF161402.8 MB0.0000
rnnt-0.6b-q8_0.ggufQ8_0903.9 MB0.0000
rnnt-0.6b-q6_k.ggufQ6_K776.3 MBnot measured
rnnt-0.6b-q5_k.ggufQ5_K705.7 MBnot measured
rnnt-0.6b-q4_k.ggufQ4_K639.2 MBnot measured

tdt-0.6b-v2

Source: nvidia/parakeet-tdt-0.6b-v2 · TDT transducer (FastConformer) · heads: TDT

FileVariantSizeWER vs NeMo
tdt-0.6b-v2-f16.gguf ← recommendedF161404.2 MB0.0000
tdt-0.6b-v2-q8_0.ggufQ8_0903.8 MB0.0000
tdt-0.6b-v2-q6_k.ggufQ6_K775.9 MBnot measured
tdt-0.6b-v2-q5_k.ggufQ5_K705.0 MBnot measured
tdt-0.6b-v2-q4_k.ggufQ4_K638.4 MBnot measured

tdt-0.6b-v3

Source: nvidia/parakeet-tdt-0.6b-v3 · TDT transducer (FastConformer) · heads: TDT

FileVariantSizeWER vs NeMo
tdt-0.6b-v3-f16.gguf ← recommendedF161441.0 MB0.0000
tdt-0.6b-v3-q8_0.ggufQ8_0940.7 MB0.0000
tdt-0.6b-v3-q6_k.ggufQ6_K812.7 MBnot measured
tdt-0.6b-v3-q5_k.ggufQ5_K741.9 MBnot measured
tdt-0.6b-v3-q4_k.ggufQ4_K675.2 MBnot measured

ctc-1.1b

Source: nvidia/parakeet-ctc-1.1b · CTC (FastConformer) · heads: CTC

FileVariantSizeWER vs NeMo
ctc-1.1b-f16.gguf ← recommendedF162395.8 MB0.0000
ctc-1.1b-q8_0.ggufQ8_01526.3 MB0.0000
ctc-1.1b-q6_k.ggufQ6_K1301.7 MBnot measured
ctc-1.1b-q5_k.ggufQ5_K1178.5 MBnot measured
ctc-1.1b-q4_k.ggufQ4_K1062.6 MBnot measured

rnnt-1.1b

Source: nvidia/parakeet-rnnt-1.1b · RNNT transducer (FastConformer) · heads: RNNT

FileVariantSizeWER vs NeMo
rnnt-1.1b-f16.gguf ← recommendedF162425.2 MB0.0000
rnnt-1.1b-q8_0.ggufQ8_01554.7 MB0.0000
rnnt-1.1b-q6_k.ggufQ6_K1331.2 MBnot measured
rnnt-1.1b-q5_k.ggufQ5_K1207.9 MBnot measured
rnnt-1.1b-q4_k.ggufQ4_K1091.9 MBnot measured

tdt-1.1b

Source: nvidia/parakeet-tdt-1.1b · TDT transducer (FastConformer) · heads: TDT

FileVariantSizeWER vs NeMo
tdt-1.1b-f16.gguf ← recommendedF162425.3 MB0.0000
tdt-1.1b-q8_0.ggufQ8_01554.8 MB0.0000
tdt-1.1b-q6_k.ggufQ6_K1331.2 MBnot measured
tdt-1.1b-q5_k.ggufQ5_K1207.9 MBnot measured
tdt-1.1b-q4_k.ggufQ4_K1091.9 MBnot measured

tdt_ctc-1.1b

Source: nvidia/parakeet-tdt_ctc-1.1b · Hybrid TDT+CTC (FastConformer) · heads: TDT + CTC

FileVariantSizeWER vs NeMo
tdt_ctc-1.1b-f16.gguf ← recommendedF162429.5 MB0.0000
tdt_ctc-1.1b-q8_0.ggufQ8_01559.0 MB0.0000
tdt_ctc-1.1b-q6_k.ggufQ6_K1335.4 MBnot measured
tdt_ctc-1.1b-q5_k.ggufQ5_K1212.1 MBnot measured
tdt_ctc-1.1b-q4_k.ggufQ4_K1096.1 MBnot measured
WER (word error rate) is computed against the upstream NeMo reference on tests/fixtures/speech.wav (LibriSpeech 2086-149220-0033, ~7.4 s, English). 0.0 = byte-for-byte identical transcript. See parity.md and quantization.md.

nemotron-3.5-asr-streaming-0.6b

Source: nvidia/nemotron-3.5-asr-streaming-0.6b · Cache-aware streaming RNNT (FastConformer, 24 encoder layers), multilingual (40 language locales), conditioned on a language prompt · heads: RNNT (streaming, prompt-conditioned) · License: OpenMDW 1.1

FileVariantSizeWER vs NeMo
nemotron-3.5-asr-streaming-0.6b-f16.gguf ← recommendedF161484.3 MB0.0000
nemotron-3.5-asr-streaming-0.6b-q8_0.ggufQ8_0983.7 MB0.0000
nemotron-3.5-asr-streaming-0.6b-q6_k.ggufQ6_K855.7 MBnot measured
nemotron-3.5-asr-streaming-0.6b-q5_k.ggufQ5_K784.8 MBnot measured
nemotron-3.5-asr-streaming-0.6b-q4_k.ggufQ4_K718.1 MBnot measured
This model runs both offline and as a cache-aware streaming model, and it is the only one here that takes a target language. Pass --lang <locale> (for example en-US, de-DE, es-ES, ja-JP), or leave it at the default auto to let the model detect the language. The WER column is agreement with NeMo's transcript on the two short test clips in the parakeet.cpp repository (speech.wav, clip.wav), not accuracy on a speech corpus. The F32 conversion was compared with NeMo for en, de, es, ja-JP and auto, offline and streaming, and matched every time. Q80 was compared on `speech.wav` in English, offline. The Q6K, Q5K and Q4K files have not been compared yet. See parity.md. The small prompt layers, the LSTM and the feature extractor stay F32 in every quantization.
bash
huggingface-cli download mudler/parakeet-cpp-gguf nemotron-3.5-asr-streaming-0.6b-f16.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/nemotron-3.5-asr-streaming-0.6b-f16.gguf --input audio.wav --lang de-DE

nemotron-3-diarization

Source: nvidia/Nemotron-3-Diarization · Speaker diarization (Sortformer, up to 8 speakers, streaming speaker cache) · License: OpenMDW 1.1

FileVariantSizeSegments vs NeMo
nemotron-3-diarization-f16.gguf ← recommendedF16200.7 MBidentical
nemotron-3-diarization-q8_0.ggufQ8_0108.7 MBidentical on 2 of 3 clips, 99.5% of frames on the third
Checked against NeMo on a 23.6 s and a 68.5 s two-speaker clip (offline and streaming: every segment identical, to the 10 ms frame) and a 12.3 min three-speaker clip (F16 100%, Q80 99.9% of speech frames). Q80 can flip a frame whose probability sits right at the 0.5 threshold, splitting a segment (99.5% of frames on a 31.5 s clip); use F16 when segment-exact output matters. Answers "who spoke when"; pair it with any ASR model above for speaker-attributed transcripts. See diarization.md.
bash
build/examples/cli/diarize models/nemotron-3-diarization-f16.gguf meeting.wav

ultra and redux (Moondream)

Source: moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 · TDT transducer (FastConformer), multilingual, with a voice-activity head (transcribe --vad) · heads: TDT · license: CC-BY-4.0

FileVariantSizeRuns on
ultra-f16.ggufF161441.9 MBany backend
ultra-q8_0.ggufQ8_0941.5 MBany backend
redux-packed.ggufpacked ternary encoder213.3 MBCPU only, offline only
redux-f16.ggufF16, dequantized1441.9 MBany backend
redux-q8_0.ggufQ8_0, dequantized941.5 MBany backend
There is no NeMo reference for these models, so there is no WER column. Each file was checked to decode the tests/fixtures/speech.wav clip to the expected sentence. See ternary.md for the measurements.
  • —Ultra has ordinary F16 weights and runs on any backend, like the other v3 files.
  • —Redux packed keeps the encoder linear layers as ternary weights (-1, 0 or +1 times a per-group scale) and runs a native CPU kernel. It does not run on GPU backends or in streaming mode, and parakeet.cpp refuses to load it there. For GPU use, take redux-f16.gguf or redux-q8_0.gguf.
  • —Redux F16 and Q8_0 are dequantized: the ternary weights were expanded to ordinary weights and then stored as F16 or Q8_0. They are not packed.
  • —Changes: these files are converted here, not trained. Nothing was trained or fine-tuned by the parakeet.cpp project. The models were trained by NVIDIA (the base) and Moondream (Ultra and Redux).
bash
huggingface-cli download mudler/parakeet-cpp-gguf redux-packed.gguf --local-dir models/
build/examples/cli/parakeet-cli transcribe --model models/redux-packed.gguf --input audio.wav
# Long audio: cut at pauses with the model's voice-activity head (offline only).
build/examples/cli/parakeet-cli transcribe --model models/ultra-q8_0.gguf --input long.wav --vad

silero-vad (Silero Team)

Source: snakers4/silero-vad v6.2.3 by the Silero Team · a small voice-activity detector, not a speech recognizer · 8 kHz and 16 kHz in one file · license: MIT

FileVariantSize
silero-vad-f32.ggufF322.2 MB
silero-vad-f16.ggufF16 weights, widened to F32 at load1.3 MB
  • —The files are converted from the official ONNX model with scripts/convert_silero_vad_to_gguf.py in parakeet.cpp. The GGUF records the source version and the ONNX sha256. They were converted here, not trained.
  • —They need a parakeet.cpp build that has the standalone VAD API. That code is not in a release yet, so check the parakeet.cpp repository before relying on it.
  • —Use it as a stand-alone detector (parakeet-cli vad --model silero-vad-f16.gguf --input audio.wav) or to cut long audio before transcribing with any model, including those that have no VAD head (parakeet-cli transcribe --model tdt-0.6b-v3-q8_0.gguf --input long.wav --vad --vad-model silero-vad-f16.gguf).
  • —Probabilities match onnxruntime to about 1e-6 (F32) and 1e-3 (F16) on the test clips.

vad heads (Moondream)

Source: the voice-activity head, the mel front end and the subsampler of moondream/parakeet-redux and moondream/parakeet-ultra by Moondream, derived from NVIDIA's nvidia/parakeet-tdt-0.6b-v3 · VAD only: these files cannot transcribe · license: CC-BY-4.0

FileParentSize
redux-vad.ggufredux (the head and subsampler are not ternary)9.9 MB
ultra-vad-q8_0.ggufultra-q8_06.0 MB
  • —Each file holds 20 tensors copied byte for byte from its parent, with no requantization, and records the parent file, its size and its sha256 in the metadata. They were cut out here, not trained.
  • —Output is identical to the full parent file: parakeet-cli vad --probabilities gave byte-identical JSON on a speech clip, a noisy clip and a 600 s talk.
  • —Against the 213 MB to 941 MB parents, load time falls from about 0.1 to 0.4 s to a few milliseconds, and memory for a 33 s clip from 0.6 to 1.1 GiB to about 0.25 GiB. Speed is the same as the parent's head.
  • —They need a parakeet.cpp build that can load a VAD-only file. That code is not in a release yet, so check the parakeet.cpp repository before relying on it. Use parakeet-cli vad --model redux-vad.gguf --input audio.wav.
  • —For a stand-alone detector Silero (above) is smaller and faster per core. The head is useful when you want its recall or already work with the Moondream models.

Bundles

A bundle is one GGUF file that holds several of the models above, so you pass one file instead of several. A bundle can hold an ASR model, a Silero VAD, speaker diarization, sound-event tagging (CED) and speaker identification. A bundle is only packaging: each model is copied byte for byte with the type it was published with, and nothing was re-quantized, trained or fine-tuned. These files are converted here, not trained. The format is described in bundle.md.

FileSizeContentsRuns on
parakeet-bundle-small.gguf337.9 MBtdtctc-110m Q80, Nemotron-3-Diarization Q80, CED-small Q80, WeSpeaker ResNet34-LM F32, Silero VAD F16any backend
parakeet-bundle-standard.gguf1100.8 MBtdt-0.6b-v3 Q80, Nemotron-3-Diarization Q80, CED-small Q8_0, WeSpeaker ResNet34-LM F32, Silero VAD F16any backend
parakeet-bundle-moondream-redux.gguf214.6 MBredux-packed (with its own VAD head), Silero VAD F16CPU only, offline only
A bundle needs a parakeet.cpp build that includes the bundle code (pull request 85). That code is on the master branch but not in a release yet (the latest release, v0.5.0, does not have it), so build from source for now. Older builds refuse a bundle with a load error. They never read the wrong weights. The bundles were checked on CPU only: the output of each component equals the output of its single-model file (the transcript, the diarization segments, the CED class scores, the speaker embeddings and the Silero probabilities). GPU backends, streaming ASR from a bundle, macOS and Windows were not tested.
bash
huggingface-cli download mudler/parakeet-cpp-gguf parakeet-bundle-small.gguf --local-dir models/

# List the components, licences and credits (reads only the header)
build/examples/cli/parakeet-cli info models/parakeet-bundle-small.gguf

# Transcribe with the ASR component; add --vad to cut long audio with the Silero component
build/examples/cli/parakeet-cli transcribe --model models/parakeet-bundle-small.gguf --input audio.wav

# Who spoke when, with the diarization component
build/examples/cli/diarize models/parakeet-bundle-small.gguf meeting.wav

# Speaker-attributed transcript with sound events: one file passed for every role
build/examples/cli/parakeet-cli scene --model models/parakeet-bundle-small.gguf \
    --diar models/parakeet-bundle-small.gguf --sound models/parakeet-bundle-small.gguf \
    --input meeting.wav
# Add --speakers models/parakeet-bundle-small.gguf --registry people.bin to name known voices

The moondream-redux bundle has the ASR and Silero components only. Use parakeet-cli transcribe --model models/parakeet-bundle-moondream-redux.gguf --input audio.wav --vad: it cuts at pauses with the Silero component, and --vad-component asr uses the Redux head instead. When a bundle has more than one component of a kind, name one with --component, --asr-component, --diar-component, --sound-component or --speakers-component.

Licences. A bundle has no single licence, so the header says other and each component keeps the licence of the model it was converted from. The credit, the licence link and the changes are in the file header (parakeet-cli info shows them), in NOTICE-parakeet-bundle-<name>.txt next to each bundle, and the full licence texts are in the `licenses/` folder of this repo. Keep these notices when you redistribute a bundle.

ComponentModelLicenceCredit
ASR (small)nvidia/parakeet-tdt_ctc-110mCC-BY-4.0NVIDIA
ASR (standard)nvidia/parakeet-tdt-0.6b-v3CC-BY-4.0NVIDIA
ASR (moondream-redux)moondream/parakeet-redux, derived from parakeet-tdt-0.6b-v3CC-BY-4.0Moondream and NVIDIA
Diarizationnvidia/Nemotron-3-DiarizationOpenMDW-1.1NVIDIA
Sound eventsmispeech/ced-smallApache-2.0 (see the note below)Heinrich Dinkel et al., Xiaomi (mispeech)
Speaker identificationWespeaker/wespeaker-voxceleb-resnet34-LMCC-BY-4.0 (see the note below)the WeSpeaker project
VADsnakers4/silero-vadMITCopyright (c) 2020-present Silero Team
  • —CED: the mispeech/ced-* model cards say Apache-2.0, and the bundle follows them. The upstream code repository is GPL-3.0 and the original checkpoint records say CC-BY-4.0, so the licence of the weights is not consistent upstream. It has not been confirmed with the authors. The CC-BY credit to the authors is kept in the meantime.
  • —WeSpeaker: the file is converted from voxceleb_resnet34_LM.onnx of Wespeaker/wespeaker-voxceleb-resnet34-LM, whose card says CC-BY-4.0. The card of the plain wespeaker-voxceleb-resnet34 says Apache-2.0, but the WeSpeaker project states in its documentation that its VoxCeleb-trained models follow CC-BY-4.0, so the bundle uses CC-BY-4.0 and credits the WeSpeaker project. The speaker models are trained on VoxCeleb. Whether a trained model is derived from its training data is a legal question that this project does not settle.
  • —Changes: the models are converted to GGUF here and, for the ASR and diarization models, quantized to Q8_0 (the redux-packed ASR component is the published packed file). Nothing was trained or fine-tuned.
  • —The end-of-utterance model and the audeering voice-analysis heads are never put in a bundle: their licences do not allow it.

Quantization notes

Quantization is applied only to the large linear weights fed directly into ggml_mul_mat (encoder FFN + attention projections, subsampling output projection, joint enc/pred projections). All other tensors (mel filterbank, LSTM prediction net, conv kernels, batch_norm stats, norms, biases, embeddings) stay F32.

Usage

bash
# 1. Clone + build parakeet.cpp
git clone https://github.com/mudler/parakeet.cpp
cd parakeet.cpp
cmake -B build -DPARAKEET_BUILD_CLI=ON && cmake --build build -j

# 2. Download one quant (F16 recommended)
huggingface-cli download mudler/parakeet-cpp-gguf tdt_ctc-110m-f16.gguf --local-dir models/

# 3. Transcribe
build/examples/cli/parakeet-cli transcribe \
    --model models/tdt_ctc-110m-f16.gguf \
    --input audio.wav

License

Licences differ by model, so the front matter says license: other. Each file family follows the licence of the model it was converted from:

The notes below add detail. ultra-*.gguf and redux-*.gguf are converted from moondream/parakeet-ultra and moondream/parakeet-redux by Moondream, which are derived from NVIDIA's parakeet-tdt-0.6b-v3. Both are also CC-BY-4.0: credit Moondream and NVIDIA when you use these files. They were converted here, not trained, and the Redux F16 and Q80 files are dequantized from the ternary weights. `redux-vad.gguf` and `ultra-vad-q80.gguf hold only the VAD head, front end and subsampler of the Moondream models, cut out of the files above under the same CC-BY-4.0 terms: credit Moondream and NVIDIA; they were cut out here, not trained. silero-vad-.gguf` is converted from [Silero VAD](https://github.com/snakers4/silero-vad) v6.2.3 and is released under the MIT license, Copyright (c) 2020-present Silero Team; it was converted here, not trained. `nemotron-3-diarization-.gguf is derived from nvidia/Nemotron-3-Diarization and nemotron-3.5-asr-streaming-0.6b-*.gguf` from nvidia/nemotron-3.5-asr-streaming-0.6b. Both are governed by the OpenMDW License Agreement, version 1.1. The parakeet.cpp runtime is MIT-licensed.