Team Ai
Datasetpublic

AIMS-RAIL/RAIL

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark NeurIPS 2026 Hongyu Jin1,*, Siyi Wang1,*, Yang Xiao1,*, Jiaheng Dong1,*, Shihong Tan4, Kaiyuan Peng1, Georgiana Juravle2, Shanquan Chen3, Gongping Huang4, Hong Jia5, Eun-Jung Holden1, James Bailey6, Ting Dang1,† 1The University of Melbourne, 2Alexandru Ioan Cuza University of Iași, 3The University of Hong Kong, 4Wuhan University, 5The University of Auckland, 6Monash University… See the full description on the dataset page: https://huggingface.co/datasets/AIMS-RAIL/RAIL.

sourceHugging Facecc-by-4.0updated 6d agoView on Hugging Face
6likes373downloads
Dataset Card

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

NeurIPS 2026

Hongyu Jin<sup>1,\</sup>, Siyi Wang<sup>1,\</sup>, Yang Xiao<sup>1,\</sup>, Jiaheng Dong<sup>1,\</sup>, Shihong Tan<sup>4</sup>, Kaiyuan Peng<sup>1</sup>, Georgiana Juravle<sup>2</sup>, Shanquan Chen<sup>3</sup>, Gongping Huang<sup>4</sup>, Hong Jia<sup>5</sup>, Eun-Jung Holden<sup>1</sup>, James Bailey<sup>6</sup>, Ting Dang<sup>1,†</sup>

<sup>1</sup>The University of Melbourne, <sup>2</sup>Alexandru Ioan Cuza University of Iași, <sup>3</sup>The University of Hong Kong, <sup>4</sup>Wuhan University, <sup>5</sup>The University of Auckland, <sup>6</sup>Monash University

<sup>\*</sup>Equal contribution. <sup>†</sup>Corresponding author.

Contact: Hongyu Jin (hongyuj1@student.unimelb.edu.au), Ting Dang (ting.dang@unimelb.edu.au)

Project page · arXiv · Paper on Hugging Face · Dataset on Hugging Face

TL;DR: RAIL evaluates large audio-language models the way cognitive science assesses human listeners. Grounded in Cattell–Horn–Carroll (CHC) theory, it organizes 5,306 audio questions into five core auditory capabilities and 32 subcapabilities, and compares 26 LALMs with a human baseline. Models do well on knowledge inherited from text pretraining, while fine-grained auditory perception and processing efficiency remain below human level.

RAIL: five CHC capabilities and 32 subcapabilities

Taxonomy and statistics

CapabilityCodeTasksSamplesTotal dur. (h)Mean dur. (s)Median dur. (s)Config
Auditory ProcessingGa711704.9815.314.0auditory_processing
ReasoningGf33223.4939.0425.04reasoning
MemoryGsm/Glr6100012.7745.9928.04memory
Processing EfficiencyGs918002.014.021.81processing_efficiency
KnowledgeGc/Gkn710147.1825.4810.49knowledge
Overall32530630.4320.658.0all

Durations are summed over all clips of an item and computed from the audio files.

Tasks and subtasks

Every task is also a config, named by its CHC code (U1/U9 → U1_U9).

CapabilityTask (code)SamplesAnswer formatSubtasks (samples)
Auditory ProcessingSpeech Sound Discrimination (US)153multiple_choiceEmotional Prosody Comparison (55); Confusable Word Identification (58); Minimal-Pair Discrimination (40)
Auditory ProcessingResistance to Auditory Stimulus Distortion (UR)181multiple_choiceBabble Noise (65); Band-Limiting (40); Background Music (36); Reverberation (40)
Auditory ProcessingMaintaining and Judging Rhythm (U8)160multiple_choiceBeat Regularity Detection (40); Meter Identification (38); Tempo Change Detection (30); Tempo Comparison (52)
Auditory ProcessingAbsolute Pitch (UP)200multiple_choiceMIDI Pitch Identification (80); Note Name Identification (50); Relative Pitch Comparison (40); Frequency Estimation (30)
Auditory ProcessingMusical Discrimination and Judgment (U1/U9)102multiple_choiceAesthetic and Emotional Judgment (8); Genre Recognition (35); Instrument Identification (50); Musical Structure Analysis (9)
Auditory ProcessingSound Localization (UL)182multiple_choiceDirection Identification (82); Azimuth Estimation (50); Distance Estimation (30); Motion Trajectory (20)
Auditory ProcessingPhonetic Coding (PC)192multiple_choiceIsolated Phoneme Identification (42); Concurrent Phoneme Segregation (50); Staggered-Onset Phoneme Identification (50); Stereo-Separated Phoneme Identification (50)
ReasoningInduction (I)68multiple_choicePattern Abstraction (34); Rule Induction (34)
ReasoningGeneral Sequential Reasoning (RG)154multiple_choiceSequential Reasoning with General Rules (124); Auditory Cognitive Puzzle (30)
ReasoningQuantitative Reasoning (RQ)100openended, multiplechoiceMath Reasoning (50); Counting (50)
MemoryMemory Span (MS)150multiple_choice—
MemoryAssociative Memory (MA)150multiple_choice—
MemoryMeaningful Memory (MM)150multiple_choiceMulti-Turn Dialogue (100); Narrated Passage (50)
MemoryFree Recall Memory (M6)200free_recall—
MemoryMemory for Sound Patterns (UM)200closedlabel, multiplechoiceN-Back (70); Speaker Tracking (70); Prosodic Contour Matching (60)
MemoryWorking Memory (WM)150multiple_choice—
Processing EfficiencyPerceptual Speed (P)200closed_label—
Processing EfficiencyRate-of-Test-Taking (R9)200closed_label—
Processing EfficiencyNumber Facility (N)200closed_label—
Processing EfficiencyReading Speed (RS)200closed_label—
Processing EfficiencySimple Reaction Time (R1)200closed_label—
Processing EfficiencyChoice Reaction Time (R2)200closed_label—
Processing EfficiencySemantic Processing Speed (R4)200closed_label—
Processing EfficiencyMental Comparison Speed (R7)200closed_label—
Processing EfficiencyInspection Time (IT)200closed_label—
KnowledgeGeneral (Verbal) Information (K0)143multiple_choice—
KnowledgeLanguage Development (LD)181multiple_choice—
KnowledgeListening Ability (LS)150multiple_choice—
KnowledgeForeign Language Proficiency (KL)193multiple_choice—
KnowledgeGeography Achievement (A5)65multiple_choice—
KnowledgeMechanical Knowledge (MK)141multiple_choiceAnomaly Detection (47); Machine Identification (47); Machine Function Inference (47)
KnowledgeKnowledge of Behavioral Content (BC)141multiple_choice—

Loading

python
from datasets import load_dataset
from huggingface_hub import snapshot_download

root = snapshot_download("AIMS-RAIL/RAIL", repo_type="dataset")  # audio + manifests
ds = load_dataset("AIMS-RAIL/RAIL", "all", split="test")        # or "memory", "UL", "U1_U9", ...

item = ds[0]
audio_files = [f"{root}/{p}" for p in item["audio_paths"]]
print(item["prompt"])

Fields

FieldDescription
idUnique item id, <capability>-<task>-<nnnn> (e.g. ga-ul-0001).
capability, capability_codeCore capability and CHC broad-ability code (Ga, Gf, Gsm/Glr, Gs, Gc/Gkn).
task, task_codeSub-capability (CHC narrow ability) and its code, as named in the paper.
subtaskFiner task design within a sub-capability; null where the task has a single design.
answer_formatmultiple_choice, closed_label (answer is one of choices, no letters shown), open_ended, or free_recall.
questionQuestion stem without options or output instructions.
choicesOption texts without letter prefixes (multiple_choice), or the allowed output labels (closed_label).
answerGold answer text. For free_recall, one target item per line.
answer_labelGold option letter for multiple_choice; null otherwise.
promptFull model-facing prompt used in the paper's evaluation (question, options and output instruction).
audio, audio_paths, num_audioPrimary clip, all clips in presentation order, and clip count (paths are relative to the repo root).
duration_sTotal audio duration of the item in seconds.
source_idItem id in the source manifest the item was generated from.
legacy_id, legacy_configId and config name in RAIL v1, for mapping earlier results.
metadata_jsonOriginal source row as JSON text (generation parameters, conditions, provenance).

Evaluation

The evaluation toolkit (Hugging Face model inference, the paper's scoring rules and an English report) is in the eval/ folder of this repository and of the GitHub repository (https://github.com/AIMS-RAIL/RAIL):

bash
cd eval
python run_rail_eval.py all --model Qwen/Qwen2-Audio-7B-Instruct --out runs/qwen2_audio

Models answer in the form Reason: ...; Answer: ... (reason capped at 20 words, except Reasoning items). Scoring follows the paper:

  • —Multiple choice / closed label: the answer must contain every token of the gold answer (or its letter) and no token unique to another option (tokenizer [A-Za-z0-9]+). An LLM-as-judge score (GPT-5.4) is reported alongside.
  • —Processing Efficiency (Gs): B-AUC, the normalised area under accuracy as a function of the reason-length budget (0–50 tokens).
  • —Free recall (M6): token recall of the target items in answer, pooled over items.
  • —Open-ended math (RQ, Math Reasoning): the final numeric answer is compared with answer.

Changes from v1

  • —Labels follow the paper: capability → task (32) → subtask; v1 benchmark/subset/task/ability are replaced.
  • —Configs are organised by capability and by task instead of by source manifest.
  • —Ids are unique (v1 had 35 duplicated ids); v1 ids are kept in legacy_id.
  • —Memory prompts and options are the exact ones used in evaluation (v1 lacked prompts for Memory Span, options for Working/Meaningful Memory prompts, and options for prosodic matching).
  • —Speaker-tracking items (Memory for Sound Patterns) use the evaluated question wording and options; v1 carried an earlier wording for the same audio.
  • —Options no longer carry letter prefixes; efficiency tasks list their allowed output labels.
  • —duration_s is filled for every item from the audio files.
  • —Audio is stored as audio/<capability>/<task>/<id>[_k].<ext>.
  • —v1 remains available: load_dataset("AIMS-RAIL/RAIL", revision="v1.0").

Citation

bibtex
@inproceedings{jin2026rail,
  title     = {RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models
               with a CHC-Grounded Benchmark},
  author    = {Jin, Hongyu and Wang, Siyi and Xiao, Yang and Dong, Jiaheng and
               Tan, Shihong and Peng, Kaiyuan and Juravle, Georgiana and Chen, Shanquan and
               Huang, Gongping and Jia, Hong and Holden, Eun-Jung and Bailey, James and
               Dang, Ting},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.11260}
}

License

CC-BY-4.0. Some items are derived from existing corpora; use them within the terms of those sources.