anonymous2026082026/MMAG-Benchmark
MMAG: A Multi‑Control Mixed Audio Generation Benchmark MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment. Dataset Structure The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific evaluation dimension:
Each subset is provided as a separate Hugging Face configuration.
File Structure
MMAG/
├── README.md
├── dataset_infos.json
├── audio_files/
│ └── *.wav
├── prompt_audio/
│ └── *.wav
├── main_set/
│ └── metadata.jsonl
├── voice_cloning_set/
│ └── metadata.jsonl
└── timestamp_set/
└── metadata.jsonlLoading the Dataset
Option 1 Soundfile
Install the required dependencies:
pip install datasets soundfile huggingface_hub Set HF_ENDPOINT
export HF_ENDPOINT="https://anonymous-hf.com"Load any subset via Hugging Face datasets:
from huggingface_hub import snapshot_download
import soundfile as sf
from datasets import load_dataset
REPO_ID = "w8zbhebrzaoq"
# Step 1: Download the entire dataset repository
snapshot_download(
repo_id=REPO_ID,
repo_type="dataset",
local_dir="./mmag_data"
)
# Step 2: Load the metadata
dataset = load_dataset(REPO_ID, "main_set", split="test")
# Step 3: Access audio files from local directory
example = dataset[0]
audio, sr = sf.read(f"./mmag_data/{example['file_name']}")
print(f"Audio shape: {audio.shape}, Sample rate: {sr}")For the voice cloning subset, the corresponding voice prompt is loaded from the voice_prompt field:
prompt_audio, sr = sf.read(f"./mmag_data/{example['voice_prompt']}")Option 2 TorchCodec
pip install datasets torchcodecfrom datasets import load_dataset, Audio
REPO_ID = "w8zbhebrzaoq"
dataset = load_dataset(REPO_ID, "main_set", split="test")
dataset = dataset.cast_column("file_name", Audio())
example = dataset[0]
audio = example["file_name"]["array"]
sr = example["file_name"]["sampling_rate"]Annotation Fields
Each subset's metadata file (metadata.jsonl) contains the following core fields:
Main Set (main_set/metadata.jsonl)
Voice Cloning Set (voicecloningset/metadata.jsonl)
Timestamp Set (timestamp_set/metadata.jsonl)
Caption Annotation Coverage
The caption field in each subset is not a simple description; it is a richly structured annotation that covers multiple acoustic and semantic dimensions. The following table summarizes the types of information encoded in the captions:
All captions are generated through a multi-expert annotation pipeline and verified by human quality control to ensure factual accuracy and consistency.
License
This dataset is released under the CC BY 4.0 license for non-commercial research use. All source audio materials are sourced from publicly available datasets and are used in accordance with their original licenses.
