datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CAIMAN-ASR-BackgroundNoise
Dataset Card for Myrtle/CAIMAN-ASR-BackgroundNoise
This dataset provides background noise audio, suitable for noise augmentation
while training Myrtle.ai's CAIMAN-ASR models.
Dataset Details
Dataset Description
Curated by: Myrtle.ai
License: Myrtle.ai's modifications to the source data are licensed under
the CC BY 4.0 license.
Some of the original data is under the CC BY 3.0 license; the rest is in the public domain.
Please see the Source Data section… See the full description on the dataset page: https://huggingface.co/datasets/Myrtle/CAIMAN-ASR-BackgroundNoise.musicai-background-music-audio-llm-benchmark
Does Background Music Matter to Speech in Pre-trained Language Models
The completed September 2026 study covers 8 model families, 55 instrumental recordings, and 10 evaluation settings. It studies how adding background music to the same spoken question changes model responses.
Latest release and artifact guide
Technical report PDF
Complete LaTeX project
LaTeX GitHub repository
Matrices, figures, and supporting data
Regenerated speech and mixtures: 550 archives / 250,800… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/musicai-background-music-audio-llm-benchmark.common_voice_22_yue_w_background_captionMerged JackyHoCL/urban-noise-uganda-61k-caption, OpenSound/AudioCaps
TODO: convert to MP3, reduce size
backgroundmusicASR-WPM-And-Background-Noise-Eval
ASR WPM and Background Noise Evaluation Dataset
A dataset of annotated audio recordings for evaluating how different factors affect Whisper (and other ASR/STT systems) transcription accuracy.
Purpose
This dataset provides controlled audio samples with annotations to evaluate ASR performance across:
Speaking pace (fast, normal, slow, mumbled, whispered, weird voices)
Background noise (cafe, music, conversations in various languages, traffic, sirens, etc.)
Microphone… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/ASR-WPM-And-Background-Noise-Eval.background-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Full version of dataset is… See the full description on the dataset page: https://huggingface.co/datasets/kakadong2018/background-noise-detection-dataset.background
Background Noise Dataset
This dataset contains 3 audio recordings of 2 different background noise classes.
Dataset Statistics
Total Audio Files: 3
Total Classes: 2
Format: WAV (AudioFolder with metadata.jsonl)
Classes and Descriptions
The dataset covers the following background noises:
Label
Description
fan_noise
Fan noise background
white_noise
White noise background
Structure
The dataset is organized in a folder structure… See the full description on the dataset page: https://huggingface.co/datasets/sdialog/background.background-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Full version of dataset is availible… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/background-noise-detection-dataset.background_noisebackground-noise-detection-dataset
Speech-Free Background Noise Dataset — Real-World, Non-Synthetic (50+ Hours)
Dataset summary
50+ hours of real-world urban environmental/ambient background noise (field recordings) without intelligible speech (speech-free), from three scenes: airport, street, subway. The dataset is non-synthetic and intended for speech enhancement via noise augmentation and sound event detection (SED) as “clean background”/negative class
Purpose and usage scenarios
Speech… See the full description on the dataset page: https://huggingface.co/datasets/NashAli/background-noise-detection-dataset.background_noisebackground-moving
