Team Ai
Datasetpublic

NathanG1en/fandub-studio-diarization-ab-test

fandub-studio diarization A/B test Test harness comparing two speaker-diarization approaches on a synthetic "fandub-like" clip (three voices + a music bed at -13 dB + noise at -30 dB), run 2026-09-20 on 2 CPU cores. Result Duration-weighted speaker attribution score (how much of the 66.2s of speech was assigned to the correct speaker, judged against ground truth): pipeline segments score A — repo spectral_vad algorithm (energy VAD + amplitude/ZCR… See the full description on the dataset page: https://huggingface.co/datasets/NathanG1en/fandub-studio-diarization-ab-test.

sourceHugging Faceupdated 20d agoView on Hugging Face
0likes61downloads
Dataset Card

fandub-studio diarization A/B test

Test harness comparing two speaker-diarization approaches on a synthetic "fandub-like" clip (three voices + a music bed at -13 dB + noise at -30 dB), run 2026-09-20 on 2 CPU cores.

Result

Duration-weighted speaker attribution score (how much of the 66.2s of speech was assigned to the correct speaker, judged against ground truth):

pipelinesegmentsscore
A — repo spectral_vad algorithm (energy VAD + amplitude/ZCR clustering)139.1%
B — Silero VAD + ECAPA speaker embeddings + clustering2586.1%

Pipeline A's energy-VAD threshold (0.35 x mean) collapses most of the mix when a music bed is present: its 13 segments cover only ~6s of the 66.2s of speech. Pipeline B clustered all three speakers with zero cross-speaker leakage. Total wall time for both pipelines on a 75s clip: 15.9s.

Files

  • —ab_test.py — the full harness (test-audio builder, both pipelines, scoring, stem rendering). Rerunnable end to end.
  • —ground_truth.rttm — true labels (speakers 374, 1487, 2277 from LibriSpeech test material)
  • —pipeline_a.rttm / pipeline_b.rttm — predicted segments per pipeline
  • —versions.json — exact package pins used

To rerun

bash
pip install "torch==2.9.1" silero-vad==6.2.2 speechbrain==1.1.1 \
  soundfile==0.14.0 scikit-learn==1.9.1 pandas huggingface_hub
python ab_test.py

Needs network on first run (LibriSpeech parquet shards + ECAPA checkpoint, ~83 MB). CPU-only; no ffmpeg required.