datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ACE-Data-0
ACE-Data-0
Human-Centric Ambient Capture as Embodied Data Engine
S-Lab, Nanyang Technological University, Singapore
·
ACE Robotics
ACE turns real home environments into spatially calibrated, temporally synchronized recording studios for embodied AI.
▶ Demo video
·
Full story, figures, and interactive examples on the blog
What this is
Learning to act in the physical… See the full description on the dataset page: https://huggingface.co/datasets/ACERobotics/ACE-Data-0.SSL_bus_compressor_control_voltage_dataset
SSL Bus Compressor Control Voltage Dataset
This repository hosts a dataset for the analysis and modeling of dynamic range compression in an SSL-style bus compressor. The dataset includes paired input/output audio signals along with the corresponding control signal applied to the voltage-controlled amplifier, expressed directly as gain reduction in decibels.
This dataset is introduced in the paper, Evaluating Dynamic Range Compressor Models Using Control-Voltage Measurements: an… See the full description on the dataset page: https://huggingface.co/datasets/bthomp23/SSL_bus_compressor_control_voltage_dataset.Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.opencs2_dataset
OpenCS2 - POV Renders
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row
in the main table is one player's perspective for one round; ten POVs per round share the same tick
clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset.satb-choral-dataset
SATB Choral Source Separation Dataset (Compressed)
This dataset contains preprocessed 4-second audio chunks from the Choral Singing Dataset (CSD) formatted for SATB (Soprano, Alto, Tenor, Bass) voice source separation tasks.
This is the compressed version with 8kHz sample rate and int8 precision for smaller file sizes.
Data Structure
Folder Structure
├── chunks/ # All individual chunk .pt files
├── quality_samples/ # Sample WAV… See the full description on the dataset page: https://huggingface.co/datasets/EwanB/satb-choral-dataset.stream-data-newQuranic-Recitation-Data
🌟 Overview
Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level.
This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.Vedavani-Dataset
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
Vedavani is the first benchmark dataset for automatic speech recognition (ASR) on Vedic Sanskrit poetry, consisting of richly annotated verses from the Rig Veda and Atharva Veda. This corpus captures the unique prosodic structure, phonetic complexity, and chanting style found in traditional Vedic recitation.
🔗 Paper: Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry (ACL 2025)📁 GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/Vedavani-Dataset.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.dataCASTLE2024
What is CASTLE?
The CASTLE dataset is a large-scale, multimodal dataset designed for advancing research in lifelogging, human activity recognition, and multimodal retrieval. It provides a rich collection of time-aligned sensor and video data for analysis and benchmarking. See the Paper (or its arXiv pre-print) for more details.
You can check our website for more details.
Characteristics
Captured over four days in a controlled environment
10 participants engaged… See the full description on the dataset page: https://huggingface.co/datasets/CASTLE-Dataset/CASTLE2024.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.Barkopedia_Dog_Sex_Classification_Dataset
📦 Dataset Description
This dataset is part of the Barkopedia Challenge: https://uta-acl2.github.io/barkopedia.html
Check training data on Hugging Face:
👉 ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset
This challenge provides a dataset of labeled dog bark audio clips:
29,345 total clips of vocalizations from 156 individual dogs across 5 breeds:
Shiba Inu
Husky
Chihuahua
German Shepherd
Pitbull
Training set: 26,895 clips
13,567 female13,328 male
Test set: 2,450… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_Dog_Sex_Classification_Dataset.asr-leaderboard-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.opencs2_dataset_wds
OpenCS2 - POV Renders WebDataset
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each
sample is one player's perspective for one round; ten POVs per round share the same tick clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that player's… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset_wds.Quranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset
Check Training Data here: ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET
Dataset Description
This dataset is for Dog Age Group Classification and contains dog bark audio clips. The data is split into training, public test, and private test sets.
Training set: 17888 audio clips.
Test set: 4920 audio clips, further divided into:
Test Public (~40%): 1966 audio clips for live leaderboard updates.
Test Private (~60%): 2954 audio clips for final evaluation.
You will… See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Barkopedia_DOG_AGE_GROUP_CLASSIFICATION_DATASET.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.Bagpiper_SFT_Data
Bagpiper SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper SFT Data is the supervised fine-tuning corpus for
Bagpiper, an open-ended audio language model
that understands and generates speech, music, environmental sound, and their
mixtures through rich textual captions and planning.
The public release has exactly two configurations:
Configuration
Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.translation_dataset
translation_dataset
Synthetic, expressive, multilingual speech for cross-lingual dubbing research.
Each example pairs a style-annotated text with generated audio that clones an English
reference voice: the voice stays the same, the language changes.
~1.9M examples in the one_speaker config
17 languages
~900 distinct reference speakers
Samples
Each sample shows the generated audio followed by the English reference voice that
conditioned it.
English… See the full description on the dataset page: https://huggingface.co/datasets/mlinmg/translation_dataset.Luhya-ASR-Data-subset-642H
Luhya ASR Data Subset 642H
Luhya speech dataset for automatic speech recognition.
numberblocks-one-voice-datasetmiimo-audio-dataset
Edge Audio Dataset
출처 및 라이선스
이 데이터셋은 여러 출처를 합친 것이고, 출처마다 라이선스가 다르다.
단일 라이선스로 배포되지 않으므로 사용 전 아래를 각각 확인해야 한다.
라벨
출처
라이선스
alarm, fire_alarm, water, scream, glass_break, bicycle, gunshot, baby_cry(울음분)
AI Hub (한국지능정보사회진흥원)
⚠️ 재배포 제한 — 아래 참고
baby_cry(일부)
ESC-50 (Piczak, 2015)
CC BY-NC 3.0 — 비상업 한정, 출처 표기 필수
baby_cry(일부)
donateacry-corpus
원 저장소 확인 필요
knock
Knocking Sound Effects With Emotional Intentions
원 배포처 확인 필요
cat_meow, dog_bark, car_horn
수집처 혼재… See the full description on the dataset page: https://huggingface.co/datasets/retrina0678/miimo-audio-dataset.Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
ViMD_Dataset
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (Main EMNLP 2024)
Introduction
This document presents the accompanying dataset for the paper titled "Multi-Dialect Vietnamese: Task, Dataset, Baseline Models, and Challenges". The dataset, referred to as the Vietnamese Multi-Dialect (ViMD) dataset, is a comprehensive resource designed to capture the linguistic diversity represented by 63 provincial dialects spoken across Vietnam. The paper is… See the full description on the dataset page: https://huggingface.co/datasets/nguyendv02/ViMD_Dataset.labeled_dataFree-Spoken-Digit-Dataset
