Team Ai
Datasetpublic

Khalilah-Shields/MSU-Benchmark

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto. Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie** Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/Khalilah-Shields/MSU-Benchmark.

sourceHugging Facecc-by-nc-4.0updated 3mo agoView on Hugging Face
0likes208downloads
Dataset Card

MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios

Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.

Zhaokai Sun\, Shuai Wang\, Zhennan Lin\, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie\\*

Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China School of Intelligent Science and Technology, Nanjing University, China Shenzhen Loop Area Institute, China Base Model, Li Auto, China

<sub>\ Equal contribution &nbsp;·&nbsp; \\* Corresponding author &nbsp;·&nbsp; 📮 zksun@mail.nwpu.edu.cn</sub>

MSU-Bench is a diagnostic benchmark for evaluating how well Large Audio-Language Models (LALMs) understand who says what, and what happens between speakers, in real multi-speaker conversations. It is organized as a two-tier framework → 5 ability dimensions → 16 sub-tasks, evaluated as four-way multiple-choice questions with diagnostically-designed distractors.

  • —📄 Paper: https://arxiv.org/abs/2606.22868
  • —🌐 Demo: https://aslp-lab.github.io/msu-bench.github.io/
  • —💻 Code & pipeline: https://github.com/ASLP-lab/MSU-Bench
⚠️ License / usage: the audio is sourced from third-party copyrighted film/TV, telephone, meeting, and podcast material. This dataset is released for non-commercial academic research only (CC-BY-NC-4.0). Do not redistribute the raw media commercially.

Dataset at a glance

Total QA items2,847 (human-reviewed subset: 2,223 with verified = true)
Sub-tasks16 across 5 ability dimensions
TiersTier 1 (Speaker Grounding & Identification): 1,884 · Tier 2 (Multi-Speaker Dialogue Reasoning): 963
LanguagesEnglish: 1,421 · Chinese: 1,426
Audio clips241 .wav segments
Format4-way multiple choice, exact-match accuracy

Scenarios (media × language):

EnglishChinese
Film / TVmovieenmoviecn
Telephonetelentelcn
Meetingmeetingenmeetingcn
Podcastpodcastenpodcastcn

Directory layout

publish-huggingface/
├── README.md                 # this dataset card
├── data/
│   └── test.jsonl            # one row per question (flat, self-contained)
├── audio/                    # 241 source .wav clips, by <scenario>/<segment>/...
├── annotations/              # per-clip speaker-segment annotations (diarization, transcript, attributes)
└── build_test_jsonl.py       # script used to (re)generate test.jsonl

Data fields (data/test.jsonl)

FieldTypeDescription
uidstringStable unique id for the question
scenariostringOne of the 8 media×language scenarios
media_typestringfilm / telephone / meeting / podcast
languagestringen / zh
tierint1 = Speaker Grounding & Identification, 2 = Multi-Speaker Dialogue Reasoning
dimensionstringAbility dimension (e.g. Speaker Identification)
taskstringSub-task name in English (e.g. Speaker Retrieval)
task_zhstringOriginal Chinese task label
levelstringlevel1 / level2
qa_lengthstringSource segment length bucket (long / short)
movie / partstringSource segment identifiers
questionstringThe question prompt
question_typestringSpeaker-referencing scheme / task variant (see below)
optionslist[string]Four options, prefixed A.–D.
answerstringCorrect option letter (A/B/C/D)
answer_textstringCorrect option text
audiostringRelative path to the source clip under audio/
annotationstringRelative path to the clip annotation under annotations/
speaker_metaobjectAcoustic-anchor context (target-speaker segments, transcript, attributes)
verifiedboolWhether the item passed human review (error-free)

Speaker-referencing schemes (question_type)

ValueMeaning
no_indexTarget specified by a raw audio snippet (acoustic anchor)
time_indexTarget specified by a time range
transcript_indexTarget specified by a quoted transcript line
speaker_indexTarget specified by order of appearance
complex_indexTarget specified by a combination of cues
reverse_retrival, reverse_count, speech_index, type_textTask-specific question variants

Usage

Load the QA table

python
from datasets import load_dataset

ds = load_dataset("<your-org>/MSU-Bench", split="test")   # reads data/test.jsonl
print(ds[0]["question"], ds[0]["options"], ds[0]["answer"])

# only the human-verified subset:
verified = ds.filter(lambda r: r["verified"])

Resolve the audio

The audio / annotation columns are repo-relative paths. Download the repo once, then open them locally:

python
from huggingface_hub import snapshot_download
import os, soundfile as sf

root = snapshot_download("<your-org>/MSU-Bench", repo_type="dataset")
row = ds[0]
wav, sr = sf.read(os.path.join(root, row["audio"]))

Score a model

For each row, prompt your model with the audio (audio), any speaker_meta acoustic anchor, the question and options, and require a single letter A/B/C/D. Compare to answer and report exact-match accuracy, optionally broken down by tier, dimension, task, question_type, and language.


Construction & quality control

Automatic generation + human review: (1) dialogue-quality filtering, (2) multi-dimensional annotation (diarization, transcription, identity, sound events, paralinguistics), (3) prompt-based QA generation across tasks and referencing schemes, (4) audio-literate human verification. The full pipeline is open-sourced in the code repo.

Citation

bibtex
@inproceedings{sun2026msubench,
  title     = {MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios},
  author    = {Sun, Zhaokai and Wang, Shuai and Lin, Zhennan and Wang, Chengyou and Gao, Dehui and Cao, Yu'ang and He, Chunjiang and Zhou, Pan and Xie, Lei},
  booktitle = {Proc. Interspeech},
  year      = {2026}
}