AIMS-RAIL/RAIL
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark NeurIPS 2026 Hongyu Jin1,*, Siyi Wang1,*, Yang Xiao1,*, Jiaheng Dong1,*, Shihong Tan4, Kaiyuan Peng1, Georgiana Juravle2, Shanquan Chen3, Gongping Huang4, Hong Jia5, Eun-Jung Holden1, James Bailey6, Ting Dang1,† 1The University of Melbourne, 2Alexandru Ioan Cuza University of Iași, 3The University of Hong Kong, 4Wuhan University, 5The University of Auckland, 6Monash University… See the full description on the dataset page: https://huggingface.co/datasets/AIMS-RAIL/RAIL.
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
NeurIPS 2026
Hongyu Jin<sup>1,\</sup>, Siyi Wang<sup>1,\</sup>, Yang Xiao<sup>1,\</sup>, Jiaheng Dong<sup>1,\</sup>, Shihong Tan<sup>4</sup>, Kaiyuan Peng<sup>1</sup>, Georgiana Juravle<sup>2</sup>, Shanquan Chen<sup>3</sup>, Gongping Huang<sup>4</sup>, Hong Jia<sup>5</sup>, Eun-Jung Holden<sup>1</sup>, James Bailey<sup>6</sup>, Ting Dang<sup>1,†</sup>
<sup>1</sup>The University of Melbourne, <sup>2</sup>Alexandru Ioan Cuza University of Iași, <sup>3</sup>The University of Hong Kong, <sup>4</sup>Wuhan University, <sup>5</sup>The University of Auckland, <sup>6</sup>Monash University
<sup>\*</sup>Equal contribution. <sup>†</sup>Corresponding author.
Contact: Hongyu Jin (hongyuj1@student.unimelb.edu.au), Ting Dang (ting.dang@unimelb.edu.au)
Project page · arXiv · Paper on Hugging Face · Dataset on Hugging Face
TL;DR: RAIL evaluates large audio-language models the way cognitive science assesses human listeners. Grounded in Cattell–Horn–Carroll (CHC) theory, it organizes 5,306 audio questions into five core auditory capabilities and 32 subcapabilities, and compares 26 LALMs with a human baseline. Models do well on knowledge inherited from text pretraining, while fine-grained auditory perception and processing efficiency remain below human level.

Taxonomy and statistics
Durations are summed over all clips of an item and computed from the audio files.
Tasks and subtasks
Every task is also a config, named by its CHC code (U1/U9 → U1_U9).
Loading
from datasets import load_dataset
from huggingface_hub import snapshot_download
root = snapshot_download("AIMS-RAIL/RAIL", repo_type="dataset") # audio + manifests
ds = load_dataset("AIMS-RAIL/RAIL", "all", split="test") # or "memory", "UL", "U1_U9", ...
item = ds[0]
audio_files = [f"{root}/{p}" for p in item["audio_paths"]]
print(item["prompt"])Fields
Evaluation
The evaluation toolkit (Hugging Face model inference, the paper's scoring rules and an English report) is in the eval/ folder of this repository and of the GitHub repository (https://github.com/AIMS-RAIL/RAIL):
cd eval
python run_rail_eval.py all --model Qwen/Qwen2-Audio-7B-Instruct --out runs/qwen2_audioModels answer in the form Reason: ...; Answer: ... (reason capped at 20 words, except Reasoning items). Scoring follows the paper:
- Multiple choice / closed label: the answer must contain every token of the gold answer (or its letter) and no token unique to another option (tokenizer
[A-Za-z0-9]+). An LLM-as-judge score (GPT-5.4) is reported alongside. - Processing Efficiency (Gs): B-AUC, the normalised area under accuracy as a function of the reason-length budget (0–50 tokens).
- Free recall (M6): token recall of the target items in
answer, pooled over items. - Open-ended math (RQ, Math Reasoning): the final numeric answer is compared with
answer.
Changes from v1
- Labels follow the paper:
capability→task(32) →subtask; v1benchmark/subset/task/abilityare replaced. - Configs are organised by capability and by task instead of by source manifest.
- Ids are unique (v1 had 35 duplicated ids); v1 ids are kept in
legacy_id. - Memory prompts and options are the exact ones used in evaluation (v1 lacked prompts for Memory Span, options for Working/Meaningful Memory prompts, and options for prosodic matching).
- Speaker-tracking items (Memory for Sound Patterns) use the evaluated question wording and options; v1 carried an earlier wording for the same audio.
- Options no longer carry letter prefixes; efficiency tasks list their allowed output labels.
duration_sis filled for every item from the audio files.- Audio is stored as
audio/<capability>/<task>/<id>[_k].<ext>. - v1 remains available:
load_dataset("AIMS-RAIL/RAIL", revision="v1.0").
Citation
@inproceedings{jin2026rail,
title = {RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models
with a CHC-Grounded Benchmark},
author = {Jin, Hongyu and Wang, Siyi and Xiao, Yang and Dong, Jiaheng and
Tan, Shihong and Peng, Kaiyuan and Juravle, Georgiana and Chen, Shanquan and
Huang, Gongping and Jia, Hong and Holden, Eun-Jung and Bailey, James and
Dang, Ting},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026},
url = {https://arxiv.org/abs/2606.11260}
}License
CC-BY-4.0. Some items are derived from existing corpora; use them within the terms of those sources.
