umd-zhou-lab/AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics Project Page | Paper | Code Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward). Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full AVQA training set, generated for process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation facet, that judge the quality of an audio-reasoning response — not just final answer correctness. The rubrics are designed to be scored Yes/No by an LLM judge that has access to the same audio.
Why
Outcome-only rewards (answer correct / wrong) cause reasoning collapse in RL: the policy learns to shortcut and stops producing meaningful reasoning. These rubrics provide a dense, process-level signal that rewards grounded, focused, domain-appropriate reasoning, preventing collapse and improving generalization.
Format
One JSON object per line:
{
"id": "avqa-train-000123",
"audio_path": "audio/train/<clip>.wav",
"question": "What happened in the audio?
Choices:
A. ...
B. ...",
"reference_answer": "B. Yacht consignment",
"rubric": {
"question_domain": "audio_reasoning/environmental_sound",
"rubrics": [
{"category": "Auditory Evidence Identification & Grounding", "criterion": "...", "weight": 0.25},
{"category": "Cross-Cue Verification & Distractor Elimination", "criterion": "...", "weight": 0.20},
{"category": "Reasoning Clarity & Flow", "criterion": "...", "weight": 0.20},
{"category": "Reasoning Focus & Efficiency", "criterion": "...", "weight": 0.15},
{"category": "Domain-Specific Audio Techniques", "criterion": "...", "weight": 0.20}
],
"reference_score": 0.4
}
}The 5 rubric categories (each appears exactly once per question)
- Auditory Evidence Identification & Grounding — identifying concrete audio cues actually present and anchoring claims to them.
- Cross-Cue Verification & Distractor Elimination — triangulating multiple cues; ruling out named distractor options by absent characteristic sounds.
- Reasoning Clarity & Flow — coherent observation → inference → conclusion structure.
- Reasoning Focus & Efficiency — alignment with the question objective; no tangential over-enumeration.
- Domain-Specific Audio Techniques — context-optimal vocabulary/methods (music-theoretic, phonetic/prosodic, temporal counting).
Properties: the 5 weights sum to 1.0; reference_score (weights the reference answer would satisfy) is kept strictly below 0.5 so a stronger response has room to outscore the reference. No rubric evaluates final-answer correctness — that is handled by a separate ground-truth reward.
Audio
This dataset contains annotations only. The audio_path field references clips from the original AVQA dataset. Download the original audio from gijs/avqa-processed and join by audio_path / id:
hf download Joysw909/AVQA --repo-type dataset --local-dir ./AVQAGeneration
- Generator: Gemini 3.1 Pro (
reasoning_effort=medium,temperature=0) - Each rubric is conditioned on the actual audio clip + question + reference answer.
- Questions are phrased for the audio-only setting (the original "video" phrasing is normalized to "audio").
Citation
If you use these rubrics, please cite this dataset and the original AVQA dataset.
@article{yu2026reinforcement,
title={Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning},
author={Yu, Fangxu and Feng, Tao and Min, Dehai and Lin, Zinan and Xu, Weijia and Xu, Michael and Yu, Philip S and Liu, Ge and Zhou, Tianyi},
journal={arXiv preprint arXiv:2608.02831},
year={2026}
}License
CC BY 4.0.
