Team Ai
Modelpublic

audio-cpp/Audio-Flamingo-GGUF

sourceHugging Faceotherupdated 5d agoView on Hugging Face
3likes5.5kdownloads
Model Card

Audio Flamingo 3 and Next GGUF

For use with audio.cpp.

These self-contained GGUFs package the model tensors, tokenizer, configuration, and audio.cpp model spec for Audio Flamingo 3 and Audio Flamingo Next. Both models support transcription, captioning, and questions about speech, music, and environmental sounds.

FileUpstream checkpointSizeMaximum audio
audio-flamingo-3-bf16.ggufAudio Flamingo 315.4 GiB10 minutes
audio-flamingo-3-q8_0.ggufAudio Flamingo 38.7 GiB10 minutes
audio-flamingo-3-q4_k.ggufAudio Flamingo 35.1 GiB10 minutes
audio-flamingo-next-bf16.ggufAudio Flamingo Next15.4 GiB30 minutes
audio-flamingo-next-q8_0.ggufAudio Flamingo Next8.7 GiB30 minutes
audio-flamingo-next-q4_k.ggufAudio Flamingo Next5.1 GiB30 minutes

Usage

bash
audiocpp_cli --task asr --family audio_flamingo \
  --model /path/to/audio-flamingo-3-bf16.gguf --backend cuda \
  --audio input.wav \
  --request-option "instruct=Transcribe the input speech." \
  --text-out response.txt --log

Select either GGUF with --model. Use a question such as Describe the instruments and tempo. for general audio understanding. Requested timestamps or speaker descriptions are generated text, not structured alignment or diarization results.

For server use, configure family audio_flamingo, task asr, and mode offline, then submit audio to /v1/audio/transcriptions with the instruction in prompt.

Correctness Reference

Official Python BF16 output is the reference. Correctness comparisons and performance measurements are separate tests.

ModelRuntime / weightsResult compared with official Python
Audio Flamingo 3Official Python BF16Reference
Audio Flamingo 3audio.cpp BF16Exact transcription and music answer
Audio Flamingo 3audio.cpp Q8_0Exact transcription; coherent but different music answer
Audio Flamingo 3audio.cpp Q4_KSame transcription words with different wrapper and punctuation; coherent but different music answer
Audio Flamingo NextOfficial Python BF16Reference
Audio Flamingo Nextaudio.cpp BF16Exact short transcription; matching words and line boundaries on the three-minute test after timestamps were removed
Audio Flamingo Nextaudio.cpp Q8_0Same short transcription words with one comma omitted
Audio Flamingo Nextaudio.cpp Q4_KExact short transcription

Audio Flamingo Next generates timestamps as text rather than structured alignment. Its BF16 three-minute response used the same words and line boundaries as Python after timestamps were removed, but the generated timestamps drifted.

CUDA Performance and Quantization

The following results use an RTX 5090 and the same decoded 10-second, 16-kHz mono speech input. Python uses the official Transformers BF16 model with SDPA. audio.cpp uses the CUDA backend and 8 CPU threads. Wall time and RTF are from the second request in an already-loaded session. Peak VRAM covers the complete process and was sampled every 100 ms.

ModelRuntime / weightsWarm wall timeRTFPeak VRAM
Audio Flamingo 3Official Python BF16292.4 ms0.029216,766 MiB
Audio Flamingo 3audio.cpp BF16283.3 ms0.028316,854 MiB
Audio Flamingo 3audio.cpp Q8_0184.9 ms0.01859,956 MiB
Audio Flamingo 3audio.cpp Q4_K151.7 ms0.01526,280 MiB
Audio Flamingo NextOfficial Python BF16249.0 ms0.024916,784 MiB
Audio Flamingo Nextaudio.cpp BF16246.3 ms0.024616,890 MiB
Audio Flamingo Nextaudio.cpp Q8_0160.9 ms0.01619,992 MiB
Audio Flamingo Nextaudio.cpp Q4_K129.9 ms0.01306,314 MiB

Quantized outputs remained coherent on the transcription and music-understanding requests. Q4K is the default package and offers the lowest tested VRAM. Q80 is the higher-precision quantized alternative. Quantization can change wording, punctuation, or factual details in open-ended answers and should not be treated as exact parity with BF16.

Audio Preprocessing

For comparisons with Python, supply the same decoded 16-kHz mono WAV to both implementations. Python's optional TorchCodec loader and librosa fallback use different downmixing and resampling paths. audio.cpp averages channels and uses SOXR when resampling is needed.

Both checkpoints use the NVIDIA OneWay Noncommercial License. The upstream license is included in this directory. Review its terms before use or redistribution.

See the audio.cpp model documentation for request options, conversion, and additional usage details.