Team Ai
Modelpublic

SpeechAntiSpoofingBenchmarks/DF_Arena_500M_V_1

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes54downloads
Model Card

DF Arena 500M — Speech Anti-Spoofing Arena results

RAPTOR universal anti-spoofing model. A wav2vec 2.0 XLS-R 300M self-supervised front-end whose per-layer hidden states are combined by learnable attention pooling (a layer-wise sigmoid gate over an attention-pooled summary), then passed through a 4-block Conformer head with a class token to a 2-way classifier. FP32, deterministic first-64600-sample (~4.04 s @ 16 kHz) window, tile-repeat if shorter (no random crop, no resampling). score = softmax(logits)[bonafide]; higher = more bona fide. Official Speech-Arena-2025/DFArena500MV1 checkpoint.

Paper: arXiv:2603.06164 · Params: 436M · Checkpoint: SpeechAntiSpoofingBenchmarks/DF_Arena_500M_V_1

Arena standing

![EER% 0 on J-SPAW_LA](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 9.01 on ArAD](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 0 on DFADD](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 2.11 on SONAR](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 9.71 on DeepVoice](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 2.63 on EmoFake_test](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 0.11 on LibriSeVoc](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 2.46 on CD-ADD](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 8.4 on ODSS](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 1.87 on InTheWild](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 4.33 on DECRO](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 8 on CFAD](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 1.19 on ASVspoof2019_LA](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 3.27 on HABLA](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 7.9 on CVoiceFake_small](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 5.78 on ASVspoof2021_LA](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 15.96 on PyAra](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 2.83 on XMAD](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 3.5 on ASVspoof2021_DF](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 13.43 on ASVspoof5](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 1.97 on ADD22_eval_31](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![EER% 7.44 on ADD2023_track12_test_r1](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![1-SRR% 3.1 on EmoSpoofTTS](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![1-SRR% 1.61 on LRLspoof](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![arena tier](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m) ![arena rank](https://huggingface.co/spaces/SpeechAntiSpoofingBenchmarks/SpeechAntiSpoofingArena?system=df-arena-500m)

Live leaderboard: DF Arena 500M on the Speech Anti-Spoofing Arena

Per-dataset results (24 datasets, mean EER 5.09%)

DatasetMetricScore
J-SPAW_LAEER0%
ArADEER9.01%
DFADDEER0%
SONAREER2.11%
DeepVoiceEER9.71%
EmoFake_testEER2.63%
LibriSeVocEER0.11%
CD-ADDEER2.46%
ODSSEER8.4%
InTheWildEER1.87%
DECROEER4.33%
CFADEER8%
ASVspoof2019_LAEER1.19%
HABLAEER3.27%
CVoiceFake_smallEER7.9%
ASVspoof2021_LAEER5.78%
PyAraEER15.96%
XMADEER2.83%
ASVspoof2021_DFEER3.5%
ASVspoof5EER13.43%
ADD22eval31EER1.97%
ADD2023track12test_r1EER7.44%
EmoSpoofTTS1-SRR3.1%
LRLspoof1-SRR1.61%

EER = Equal Error Rate (lower better). 1-SRR = spoof-only complement of the Spoof Recall Rate at the model's own DeepVoice EER operating point (lower better). All rows scoring-verified (`reproduce --scoring`, Δ 0.0) and computed with the TensorRT engine (parity-verified vs PyTorch).

Usage

python
from transformers import pipeline
import librosa
pipe = pipeline("antispoofing", model="SpeechAntiSpoofingBenchmarks/DF_Arena_500M_V_1", trust_remote_code=True, device="cuda")
audio, sr = librosa.load("sample.wav", sr=16000)
print(pipe(audio))   # {'label': 'bonafide'|'spoof', 'all_scores': {...}}

Citation

bibtex
@misc{kulkarni2026compactsslbackbonesmatter,
  title={Do Compact SSL Backbones Matter for Audio Deepfake Detection? A Controlled Study with RAPTOR},
  author={Ajinkya Kulkarni and Sandipana Dowerah and Atharva Kulkarni and Tanel Alumäe and Mathew Magimai Doss},
  year={2026},
  eprint={2603.06164},
  archivePrefix={arXiv},
  primaryClass={cs.SD},
  url={https://arxiv.org/abs/2603.06164}
}