NathanG1en/fandub-studio-diarization-ab-test
fandub-studio diarization A/B test Test harness comparing two speaker-diarization approaches on a synthetic "fandub-like" clip (three voices + a music bed at -13 dB + noise at -30 dB), run 2026-09-20 on 2 CPU cores. Result Duration-weighted speaker attribution score (how much of the 66.2s of speech was assigned to the correct speaker, judged against ground truth): pipeline segments score A — repo spectral_vad algorithm (energy VAD + amplitude/ZCR… See the full description on the dataset page: https://huggingface.co/datasets/NathanG1en/fandub-studio-diarization-ab-test.
fandub-studio diarization A/B test
Test harness comparing two speaker-diarization approaches on a synthetic "fandub-like" clip (three voices + a music bed at -13 dB + noise at -30 dB), run 2026-09-20 on 2 CPU cores.
Result
Duration-weighted speaker attribution score (how much of the 66.2s of speech was assigned to the correct speaker, judged against ground truth):
Pipeline A's energy-VAD threshold (0.35 x mean) collapses most of the mix when a music bed is present: its 13 segments cover only ~6s of the 66.2s of speech. Pipeline B clustered all three speakers with zero cross-speaker leakage. Total wall time for both pipelines on a 75s clip: 15.9s.
Files
ab_test.py— the full harness (test-audio builder, both pipelines, scoring, stem rendering). Rerunnable end to end.ground_truth.rttm— true labels (speakers 374, 1487, 2277 from LibriSpeech test material)pipeline_a.rttm/pipeline_b.rttm— predicted segments per pipelineversions.json— exact package pins used
To rerun
pip install "torch==2.9.1" silero-vad==6.2.2 speechbrain==1.1.1 \
soundfile==0.14.0 scikit-learn==1.9.1 pandas huggingface_hub
python ab_test.pyNeeds network on first run (LibriSpeech parquet shards + ECAPA checkpoint, ~83 MB). CPU-only; no ffmpeg required.
