Team Ai
Datasetpublic

anonymous2026082026/MMAG-Benchmark

MMAG: A Multi‑Control Mixed Audio Generation Benchmark MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment. Dataset Structure The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026082026/MMAG-Benchmark.

sourceHugging Facecc-by-4.0updated 14d agoView on Hugging Face
0likes261downloads
Dataset Card

MMAG: A Multi‑Control Mixed Audio Generation Benchmark

MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.

Dataset Structure

The dataset is organized into three subsets, each targeting a specific evaluation dimension:

Config NameClipsDescription
main_set~4,000Primary evaluation set for holistic generation quality
voicecloningset~690Speaker identity preservation from short voice prompts
timestamp_set~1,800Temporal alignment and event timing control

Each subset is provided as a separate Hugging Face configuration.

File Structure

MMAG/
├── README.md
├── dataset_infos.json
├── audio_files/
│   └── *.wav
├── prompt_audio/
│   └── *.wav
├── main_set/
│   └── metadata.jsonl
├── voice_cloning_set/
│   └── metadata.jsonl
└── timestamp_set/
    └── metadata.jsonl

Loading the Dataset

Option 1 Soundfile

Install the required dependencies:

bash
pip install datasets soundfile huggingface_hub  

Set HF_ENDPOINT

bash
export HF_ENDPOINT="https://anonymous-hf.com"

Load any subset via Hugging Face datasets:

python
from huggingface_hub import snapshot_download
import soundfile as sf
from datasets import load_dataset
REPO_ID = "w8zbhebrzaoq"
# Step 1: Download the entire dataset repository
snapshot_download(
    repo_id=REPO_ID,
    repo_type="dataset",
    local_dir="./mmag_data"
)

# Step 2: Load the metadata
dataset = load_dataset(REPO_ID, "main_set", split="test")

# Step 3: Access audio files from local directory
example = dataset[0]
audio, sr = sf.read(f"./mmag_data/{example['file_name']}")
print(f"Audio shape: {audio.shape}, Sample rate: {sr}")

For the voice cloning subset, the corresponding voice prompt is loaded from the voice_prompt field:

python
prompt_audio, sr = sf.read(f"./mmag_data/{example['voice_prompt']}")

Option 2 TorchCodec

bash
pip install datasets torchcodec
python
from datasets import load_dataset, Audio
REPO_ID = "w8zbhebrzaoq"

dataset = load_dataset(REPO_ID, "main_set", split="test")
dataset = dataset.cast_column("file_name", Audio())

example = dataset[0]
audio = example["file_name"]["array"]
sr = example["file_name"]["sampling_rate"]

Annotation Fields

Each subset's metadata file (metadata.jsonl) contains the following core fields:

Main Set (main_set/metadata.jsonl)

FieldDescription
file_nameAudio file path (relative to root)
idUnique identifier for the clip
captionOverall caption describing the full acoustic scene (see Caption Annotation Coverage below)

Voice Cloning Set (voicecloningset/metadata.jsonl)

FieldDescription
file_nameTarget audio file path (relative to root)
voice_promptVoice prompt file path (relative to root)
captionOverall caption
idUnique identifier for the clip

Timestamp Set (timestamp_set/metadata.jsonl)

FieldDescription
file_nameAudio file path (relative to root)
idUnique identifier for the clip
captionOverall caption
timestamped_captionCaption with precise temporal boundaries

Caption Annotation Coverage

The caption field in each subset is not a simple description; it is a richly structured annotation that covers multiple acoustic and semantic dimensions. The following table summarizes the types of information encoded in the captions:

CategoryAttributes Covered in Caption
SpeechTranscription, speaker gender, age, accent, emotion
MusicInstrumentation, genre, mood
Sound EventsForeground and background sound description, ambient acoustic context
TemporalRelative ordering of events (e.g., "before", "after", "throughout"), fine-grained timestamps

All captions are generated through a multi-expert annotation pipeline and verified by human quality control to ensure factual accuracy and consistency.

License

This dataset is released under the CC BY 4.0 license for non-commercial research use. All source audio materials are sourced from publicly available datasets and are used in accordance with their original licenses.