SEAR-benchmark/SEAR
SEAR: Spoofing Evidence-Grounded Audio Reasoning SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation. SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.
SEAR: Spoofing Evidence-Grounded Audio Reasoning
SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation.
SEAR contains four complementary tasks covering acoustic evidence identification and quantification, forensic rationale grounding, and deepfake detection.
Tasks
T3 contains one question for each applicable reference finding. If no selected feature is anomalous, the audio record is retained with an empty qa_pairs list and task_status.T3 = "not_applicable"; no artificial measurement question is introduced. T4 asks the model to summarize signal-level acoustic evidence in natural language. Its reference rationale verbalizes the verified feature names, measured values, and deviation directions without revealing the authenticity label or attack identity.
Scale
The complete benchmark therefore contains 1,897,881 AQA items over 223,382 audio recordings.
Data format
Each JSONL row represents one audio recording. Questions for the same recording remain grouped in qa_pairs, allowing an ALM runner to load the audio once and answer all applicable questions.
{
"partition": "19LA_dev",
"file_id": "LA_D_1000752",
"audio_path": "data/LA/ASVspoof2019_LA_dev/flac/LA_D_1000752.flac",
"label_offline": "spoof",
"attack_metadata_offline": "A05",
"qa_pairs": [
{
"qid": "LA_D_1000752_T2",
"task": "forgery_cue_identification",
"new_task": "T2",
"question_type": "multiple_choice",
"question": "Which acoustic finding best matches the deterministic analysis of this audio?",
"options": {"A": "...", "B": "...", "C": "...", "D": "..."},
"answer": "C",
"answer_text": "...",
"metadata": {"target_feature": "voiced_ratio", "no_anomaly_target": false}
}
]
}The paths are repository-relative identifiers for alignment with locally obtained source audio. This repository does not redistribute ASVspoof audio.
Loading
Replace YOUR_NAMESPACE/SEAR with this dataset repository's identifier:
from datasets import load_dataset
t1 = load_dataset("YOUR_NAMESPACE/SEAR", "t1")
t2 = load_dataset("YOUR_NAMESPACE/SEAR", "t2")
t3 = load_dataset("YOUR_NAMESPACE/SEAR", "t3")
t4 = load_dataset("YOUR_NAMESPACE/SEAR", "t4")
example = t2["19LA_eval"][0]
print(example["audio_path"], example["qa_pairs"])Users must obtain ASVspoof 2019 LA and ASVspoof 2021 LA from their official distribution channels and resolve audio_path relative to their local copy.
Evaluation protocol and leakage prevention
label_offline, attack_metadata_offline, answer, and answer_text are evaluation-side fields. They must not be included in the model prompt. In particular:
- do not reveal authenticity labels or attack identities during inference;
- do not derive normalization statistics from development or evaluation data;
- use bona-fide training speech only when constructing a deployable reference prior;
- use development data only for model selection or threshold selection;
- reserve evaluation partitions for final reporting.
The included reference evidence is an offline benchmark target constructed with dataset labels and available attack metadata. It is oracle supervision for evaluation, not the output of a deployable detector.
Construction and validation
SEAR uses 35 utterance-level descriptors spanning cepstral, spectral, prosodic, temporal, and energy characteristics. Candidate evidence was statistically screened globally and, when annotations were available, attack-wise. T2 presents a verified finding among controlled distractors and a neutral no-anomaly option. T3 masks the measured value of a specified finding and asks the model to recover it from four candidates.
All 223,382 audio records passed automated checks for partition and audio-path alignment, task cardinality, answer mapping, option uniqueness, feature/value traceability, value masking, and exclusion of attack identifiers from model-visible text. File counts and SHA-256 checksums are provided in manifest.json.
For T4, references are generated solely from the verified structured findings and checked for feature coverage, exact measured values, deviation directions, sentence count, and forbidden class or attack inferences before release.
Intended use and limitations
SEAR is intended for evaluating evidence-grounded acoustic reasoning in ALMs and for controlled research on audio deepfake detection. The reference findings are statistical descriptors rather than exhaustive causal explanations of spoofing. Performance may vary under unseen recording conditions, languages, codecs, attacks, and demographic groups. The source benchmarks and their collection processes may introduce additional biases.
Source data and terms
This release provides AQA annotations and source-audio identifiers only. Users are responsible for obtaining the corresponding ASVspoof data and complying with its terms of use. No license is asserted here over the underlying audio recordings.
Citation
The paper citation will be added upon publication. Until then, please cite the SEAR dataset repository and the corresponding ASVspoof source benchmarks.
