Khalilah-Shields/MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun\, Shuai Wang\, Zhennan Lin\, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie\\*
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing University, China Shenzhen Loop Area Institute, China Base Model, Li Auto, China
<sub>\ Equal contribution · \\* Corresponding author · 📮 zksun@mail.nwpu.edu.cn</sub>
MSU-Bench is a diagnostic benchmark for evaluating how well Large Audio-Language Models (LALMs) understand who says what, and what happens between speakers, in real multi-speaker conversations. It is organized as a two-tier framework → 5 ability dimensions → 16 sub-tasks, evaluated as four-way multiple-choice questions with diagnostically-designed distractors.
- 📄 Paper: https://arxiv.org/abs/2606.22868
- 🌐 Demo: https://aslp-lab.github.io/msu-bench.github.io/
- 💻 Code & pipeline: https://github.com/ASLP-lab/MSU-Bench
⚠️ License / usage: the audio is sourced from third-party copyrighted film/TV, telephone, meeting, and podcast material. This dataset is released for non-commercial academic research only (CC-BY-NC-4.0). Do not redistribute the raw media commercially.
Dataset at a glance
Scenarios (media × language):
Directory layout
publish-huggingface/
├── README.md # this dataset card
├── data/
│ └── test.jsonl # one row per question (flat, self-contained)
├── audio/ # 241 source .wav clips, by <scenario>/<segment>/...
├── annotations/ # per-clip speaker-segment annotations (diarization, transcript, attributes)
└── build_test_jsonl.py # script used to (re)generate test.jsonlData fields (data/test.jsonl)
Speaker-referencing schemes (question_type)
Usage
Load the QA table
from datasets import load_dataset
ds = load_dataset("<your-org>/MSU-Bench", split="test") # reads data/test.jsonl
print(ds[0]["question"], ds[0]["options"], ds[0]["answer"])
# only the human-verified subset:
verified = ds.filter(lambda r: r["verified"])Resolve the audio
The audio / annotation columns are repo-relative paths. Download the repo once, then open them locally:
from huggingface_hub import snapshot_download
import os, soundfile as sf
root = snapshot_download("<your-org>/MSU-Bench", repo_type="dataset")
row = ds[0]
wav, sr = sf.read(os.path.join(root, row["audio"]))Score a model
For each row, prompt your model with the audio (audio), any speaker_meta acoustic anchor, the question and options, and require a single letter A/B/C/D. Compare to answer and report exact-match accuracy, optionally broken down by tier, dimension, task, question_type, and language.
Construction & quality control
Automatic generation + human review: (1) dialogue-quality filtering, (2) multi-dimensional annotation (diarization, transcription, identity, sound events, paralinguistics), (3) prompt-based QA generation across tasks and referencing schemes, (4) audio-literate human verification. The full pipeline is open-sourced in the code repo.
Citation
@inproceedings{sun2026msubench,
title = {MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios},
author = {Sun, Zhaokai and Wang, Shuai and Lin, Zhennan and Wang, Chengyou and Gao, Dehui and Cao, Yu'ang and He, Chunjiang and Zhou, Pan and Xie, Lei},
booktitle = {Proc. Interspeech},
year = {2026}
}