Team Ai
Modelpublic

audio-cpp/SAM-Audio-GGUF

sourceHugging Faceotherupdated 9d agoView on Hugging Face
1likes44kdownloads
Model Card

SAM Audio GGUF

GGUF weights for native inference with audio.cpp. SAM Audio separates a sound described by text, a masked image or video, or positive and negative time spans from an input recording.

Upstream

These packages use the original Meta checkpoints, not the separate -tv variants.

VariantGGUFPinned checkpoint
Smallsam-audio-small-f32.gguffacebook/sam-audio-small
Basesam-audio-base-f32.gguffacebook/sam-audio-base
Largesam-audio-large-f32.gguffacebook/sam-audio-large

The official SAM Audio implementation is pinned to bb4c699.

Each GGUF contains SAM weights, T5 text encoder, tokenizer, configuration, and audio.cpp model spec. No separate text encoder is needed. F32 packages retain the original weights, checked byte-for-byte against the source tensors. BF16 and Q80 packages are also available for each variant; both were converted directly from the source safetensors. Biases, normalization vectors, and scale/shift tables remain F32. Q80 denotes mixed quantized storage, not that every tensor is quantized.

Usage

bash
audiocpp_cli --task s2s --family sam_audio \
  --model /path/to/SAM-Audio-GGUF/sam-audio-small-f32.gguf --backend cuda \
  --audio input.wav --text "man speaking" --seed -1 \
  --out-dir outputs/separation --log

Outputs are target.wav and residual.wav, mono 48 kHz. Optional temporal prompts use --request-option 'anchors=[["+",0.5,2.0],["-",3.0,4.0]]'. The runtime also accepts a masked static image with --request-option reference_image_path=masked.png, or a masked video with --request-option reference_video_path=masked.mp4. Video decoding requires FFmpeg/libav runtime libraries. Use one visual option at a time and keep the video and input audio time origins aligned. Automatic span prediction and candidate ranking are not included.

See the model usage guide for all request options. Code tested on CUDA/Vulkan/Metal.

[!WARNING] All parity, performance, and VRAM results below were measured with watermarking enabled to match the official Python implementation, which adds a watermark to its output. The final audio.cpp version will return audio WITHOUT the added watermark, so these results do not describe the final implementation.

Controlled Parity With Python

Parity was tested separately from performance on CUDA with F32 weights and TF32 disabled in both implementations (NVIDIA_TF32_OVERRIDE=0). Python used seed 42; C++ replayed its exact diffusion noise and watermark bits. Both used the same 14.975-second office recording, man speaking prompt, and 16 midpoint steps. C++ computed its own encoder features, diffusion, and decoded audio.

Each row compares the final C++ waveforms directly against official Python, including bounded-memory mode. Higher cosine similarity is better; 1 is exact directional agreement, not a claim of byte-identical waveforms.

VariantC++ modeTarget cosineResidual cosine
Small F32Normal0.999999450.99999864
Small F32Bounded memory0.999999680.99999896
Base F32Normal0.999999560.99999558
Base F32Bounded memory0.999999500.99999753
Large F32Normal0.999999800.99999888
Large F32Bounded memory0.999999840.99999936

All rows exceeded 0.9999 for both outputs. Repeating each controlled C++ run within the same session produced byte-identical outputs.

Ordinary CUDA Performance With Python

Measured on an NVIDIA RTX 5090 (32 GB), using a 14.975-second recording from the official office example and the text prompt man speaking.

  • —Both implementations use the original weights, 16 midpoint steps, one candidate, and no visual or span prompts. Optional ranking and span-prediction models are not loaded in Python.
  • —Ordinary inference: no fixed or replayed random noise, no TF32 override, and no forced precision/autocast changes. C++ uses --seed -1; Python does not set a seed. Python uses PyTorch 2.11.0+cu128; C++ uses a debug build and 8 threads.
  • —Warm inference time is the median of the last three of five sequential requests in one session. It excludes model loading and file I/O. RTF is inference time divided by input duration; lower is faster.
  • —Peak VRAM is process memory sampled with nvidia-smi every 50 ms across loading and all five requests, using the same method for Python and C++.
VariantC++ warm timeC++ RTFC++ peak VRAMPython warm timePython RTFPython peak VRAM
Small F32684 ms0.04568,500 MiB746 ms0.049810,972 MiB
Base F32918 ms0.061311,010 MiB1,105 ms0.073813,414 MiB
Large F321,542 ms0.103017,826 MiB2,103 ms0.140520,568 MiB

These are text-conditioned, single-recording measurements, not a benchmark of visual prompting, automatic span prediction, or candidate ranking.

Outputs from these independent-randomness performance runs are not used for parity comparisons.

BF16 and Q8 Compared With C++ F32

The following are a separate C++-only comparison on the same RTX 5090 and 14.975-second input. Ordinary performance uses the five-request methodology above, --seed -1, and no TF32 or precision overrides. F32 was remeasured for this comparison; no Python model was run.

VariantStorageWarm timeRTFPeak VRAM
SmallF32700 ms0.046738,500 MiB
SmallBF16674 ms0.045017,366 MiB
SmallQ8_0673 ms0.044936,834 MiB
BaseF32960 ms0.0640811,010 MiB
BaseBF16785 ms0.052418,644 MiB
BaseQ8_0794 ms0.053027,524 MiB
LargeF321,566 ms0.1045717,826 MiB
LargeBF161,146 ms0.0765212,092 MiB
LargeQ8_01,051 ms0.070189,378 MiB

Output drift was measured in separate controlled C++ runs, using seed 42 and NVIDIA_TF32_OVERRIDE=0 for every dtype. Each output is compared against that variant's C++ F32 waveform, not Python. These runs supply no performance or VRAM figures.

VariantStorageTarget cosine vs F32Residual cosine vs F32
SmallBF160.9999400.999763
SmallQ8_00.9994760.999095
BaseBF160.9999180.999788
BaseQ8_00.9996230.999069
LargeBF160.9997430.999712
LargeQ8_00.9995900.999273

BF16 reduced peak VRAM by 13-32%; Q8_0 reduced it by 20-47% in this workload. Drift is expected. These short text-conditioned checks do not establish equal quality for every recording or validate reduced-precision visual prompting and bounded-memory long-form execution. F32 remains available as the reference.

Bounded-Memory Mode

Enable the experimental session option to reduce long-recording GPU workspace:

bash
--session-option sam_audio.memory_bounded=true

The option defaults to false, and its name is provisional. The existing Small GGUF predates this option; add --model-spec-override model_specs/sam_audio.json from a checkout that includes bounded-memory support. The Base and Large GGUFs include the option in their embedded specs.

The mode retains full-recording diffusion and attention. Codec convolutions use tiles with receptive-field context, and watermark LSTM state carries between tiles. It does not separate independent audio chunks and crossfade them. Intermediate sequences reside in host RAM, which still grows with input length. Transfers and overlap computation can make it slower on short inputs; small numerical differences are possible. The configured model context limit remains 400 seconds.

The following CUDA tests used a 180-second input made by repeating the same example to 180 seconds, with the same prompt and inference settings as above. Each C++ result is the first inference in a fresh session, including graph setup but excluding model loading and file I/O. Peak VRAM includes loading and inference. These are not warm medians.

VariantC++ bounded timeC++ RTFC++ peak VRAMOfficial Python, same input
Small F328.768 s0.04876,550 MiBCUDA out of memory
Base F3211.109 s0.06179,212 MiBCUDA out of memory
Large F3217.038 s0.094716,244 MiBCUDA out of memory

All three C++ runs produced the complete target and residual recordings. Python ran out of memory in its audio encoder on the same 32 GB GPU, using its default precision and allocator settings. This demonstrates completion and memory usage for this workload, not 180-second parity against Python: Python did not produce a full-length reference. Other hardware or Python memory-saving configurations may behave differently.

License

SAM Audio weights are distributed under the SAM License, dated November 19, 2025. The included T5 encoder is from Google T5 Base and retains its Apache-2.0 license. Conversion does not replace the upstream license terms.