Team Ai
Datasetpublic

SEAR-benchmark/SEAR

SEAR: Spoofing Evidence-Grounded Audio Reasoning SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation. SEAR contains four complementary tasks covering acoustic… See the full description on the dataset page: https://huggingface.co/datasets/SEAR-benchmark/SEAR.

sourceHugging Faceupdated 21d agoView on Hugging Face
1likes181downloads
Dataset Card

SEAR: Spoofing Evidence-Grounded Audio Reasoning

SEAR is an audio question-answering benchmark for testing whether audio language models can identify and quantify signal-level acoustic anomalies and use them as evidence for audio deepfake detection. Instead of evaluating only a final authenticity verdict, SEAR separates deepfake detection, forgery-cue identification, acoustic measurement, and forensic rationale generation.

SEAR contains four complementary tasks covering acoustic evidence identification and quantification, forensic rationale grounding, and deepfake detection.

Tasks

ConfigTaskFormatEvaluation target
t1Binary classificationOne multiple-choice question per audioDeepfake detection
t2Forgery cue identificationOne multiple-choice question per audioAcoustic evidence identification
t3Acoustic feature measurementZero or more multiple-choice questions per audioAcoustic evidence quantification
t4Forensic rationale generationOne open-ended question per audioForensic rationale grounding

T3 contains one question for each applicable reference finding. If no selected feature is anomalous, the audio record is retained with an empty qa_pairs list and task_status.T3 = "not_applicable"; no artificial measurement question is introduced. T4 asks the model to summarize signal-level acoustic evidence in natural language. Its reference rationale verbalizes the verified feature names, measured values, and deviation directions without revealing the authenticity label or attack identity.

Scale

Source partitionAudioT1T2T3 questionsT4
ASVspoof 2019 LA train25,38025,38025,38096,31825,380
ASVspoof 2019 LA dev24,84424,84424,844138,54924,844
ASVspoof 2019 LA eval71,23771,23771,237639,20471,237
ASVspoof 2021 LA eval101,921101,921101,921353,664101,921
Total223,382223,382223,3821,227,735223,382

The complete benchmark therefore contains 1,897,881 AQA items over 223,382 audio recordings.

Data format

Each JSONL row represents one audio recording. Questions for the same recording remain grouped in qa_pairs, allowing an ALM runner to load the audio once and answer all applicable questions.

json
{
  "partition": "19LA_dev",
  "file_id": "LA_D_1000752",
  "audio_path": "data/LA/ASVspoof2019_LA_dev/flac/LA_D_1000752.flac",
  "label_offline": "spoof",
  "attack_metadata_offline": "A05",
  "qa_pairs": [
    {
      "qid": "LA_D_1000752_T2",
      "task": "forgery_cue_identification",
      "new_task": "T2",
      "question_type": "multiple_choice",
      "question": "Which acoustic finding best matches the deterministic analysis of this audio?",
      "options": {"A": "...", "B": "...", "C": "...", "D": "..."},
      "answer": "C",
      "answer_text": "...",
      "metadata": {"target_feature": "voiced_ratio", "no_anomaly_target": false}
    }
  ]
}

The paths are repository-relative identifiers for alignment with locally obtained source audio. This repository does not redistribute ASVspoof audio.

Loading

Replace YOUR_NAMESPACE/SEAR with this dataset repository's identifier:

python
from datasets import load_dataset

t1 = load_dataset("YOUR_NAMESPACE/SEAR", "t1")
t2 = load_dataset("YOUR_NAMESPACE/SEAR", "t2")
t3 = load_dataset("YOUR_NAMESPACE/SEAR", "t3")
t4 = load_dataset("YOUR_NAMESPACE/SEAR", "t4")

example = t2["19LA_eval"][0]
print(example["audio_path"], example["qa_pairs"])

Users must obtain ASVspoof 2019 LA and ASVspoof 2021 LA from their official distribution channels and resolve audio_path relative to their local copy.

Evaluation protocol and leakage prevention

label_offline, attack_metadata_offline, answer, and answer_text are evaluation-side fields. They must not be included in the model prompt. In particular:

  • —do not reveal authenticity labels or attack identities during inference;
  • —do not derive normalization statistics from development or evaluation data;
  • —use bona-fide training speech only when constructing a deployable reference prior;
  • —use development data only for model selection or threshold selection;
  • —reserve evaluation partitions for final reporting.

The included reference evidence is an offline benchmark target constructed with dataset labels and available attack metadata. It is oracle supervision for evaluation, not the output of a deployable detector.

Construction and validation

SEAR uses 35 utterance-level descriptors spanning cepstral, spectral, prosodic, temporal, and energy characteristics. Candidate evidence was statistically screened globally and, when annotations were available, attack-wise. T2 presents a verified finding among controlled distractors and a neutral no-anomaly option. T3 masks the measured value of a specified finding and asks the model to recover it from four candidates.

All 223,382 audio records passed automated checks for partition and audio-path alignment, task cardinality, answer mapping, option uniqueness, feature/value traceability, value masking, and exclusion of attack identifiers from model-visible text. File counts and SHA-256 checksums are provided in manifest.json.

For T4, references are generated solely from the verified structured findings and checked for feature coverage, exact measured values, deviation directions, sentence count, and forbidden class or attack inferences before release.

Intended use and limitations

SEAR is intended for evaluating evidence-grounded acoustic reasoning in ALMs and for controlled research on audio deepfake detection. The reference findings are statistical descriptors rather than exhaustive causal explanations of spoofing. Performance may vary under unseen recording conditions, languages, codecs, attacks, and demographic groups. The source benchmarks and their collection processes may introduce additional biases.

Source data and terms

This release provides AQA annotations and source-audio identifiers only. Users are responsible for obtaining the corresponding ASVspoof data and complying with its terms of use. No license is asserted here over the underlying audio recordings.

Citation

The paper citation will be added upon publication. Until then, please cite the SEAR dataset repository and the corresponding ASVspoof source benchmarks.