datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Quranic-Recitation-Data
🌟 Overview
Quranic Recitation Dataset (Word-by-Word Sync) is a highly optimized, production-ready dataset containing high-quality audio recitations of the Holy Quran synchronized at the word-by-word level.
This dataset features 135 world-renowned reciters, with every Surah (114 chapters) mapped precisely to millisecond-accurate word timestamps. It is designed for modern Islamic mobile and web applications — served via a Cloudflare Edge CDN with native… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Recitation-Data.short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.Vedavani-Dataset
Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry
Vedavani is the first benchmark dataset for automatic speech recognition (ASR) on Vedic Sanskrit poetry, consisting of richly annotated verses from the Rig Veda and Atharva Veda. This corpus captures the unique prosodic structure, phonetic complexity, and chanting style found in traditional Vedic recitation.
🔗 Paper: Vedavani: A Benchmark Corpus for ASR on Vedic Sanskrit Poetry (ACL 2025)📁 GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/sanganaka/Vedavani-Dataset.LEMAS-Dataset-train
Overview
This dataset is part of LEMAS-Project (lemas-project.github.io/LEMAS-Project).
It contains a large-scale training set (150k+ hours) and a curated evaluation set
(500 utterances per language) covering 10 languages, all with word-level alignment.
Fields
key: unique utterance identifier; the first two characters indicate the language ID
audio: relative path to the MP3 audio file (in the eval set, this key is renamed to "file_name" for compatibility with the viewer)… See the full description on the dataset page: https://huggingface.co/datasets/LEMAS-Project/LEMAS-Dataset-train.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.asr-leaderboard-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the MCV… See the full description on the dataset page: https://huggingface.co/datasets/nithinraok/asr-leaderboard-datasets.Quranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.spanish-slang-stt-data
Spanish Regional Speech-to-Text Dataset
A multilingual Spanish speech recognition dataset covering 4 regional dialects for fine-tuning Whisper and other ASR models.
Dataset Description
This dataset contains ~39,000 audio samples with transcriptions across 4 Spanish-speaking regions:
Region
Samples
Description
Mexico
17,725
Mexican Spanish including CIEMPIESS corpus
Spain
11,360
Castilian Spanish from TEDx and Common Voice
Argentina
5,839
Rioplatense Spanish… See the full description on the dataset page: https://huggingface.co/datasets/shraavb/spanish-slang-stt-data.Bagpiper_SFT_Data
Bagpiper SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper SFT Data is the supervised fine-tuning corpus for
Bagpiper, an open-ended audio language model
that understands and generates speech, music, environmental sound, and their
mixtures through rich textual captions and planning.
The public release has exactly two configurations:
Configuration
Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.translation_dataset
translation_dataset
Synthetic, expressive, multilingual speech for cross-lingual dubbing research.
Each example pairs a style-annotated text with generated audio that clones an English
reference voice: the voice stays the same, the language changes.
~1.9M examples in the one_speaker config
17 languages
~900 distinct reference speakers
Samples
Each sample shows the generated audio followed by the English reference voice that
conditioned it.
English… See the full description on the dataset page: https://huggingface.co/datasets/mlinmg/translation_dataset.Luhya-ASR-Data-subset-642H
Luhya ASR Data Subset 642H
Luhya speech dataset for automatic speech recognition.
TTS_DATASET
Phonemized-VCTK (speech + features)
Phonemized-VCTK is a light-repack of the VCTK corpus that bundles—per utterance—
the raw audio (wav/)
the plain transcript (txt/)
the IPA phoneme string (phonemized/)
frame-level pitch-aligned segments (segments/)
sentence-level context embeddings (context_embeddings/)
speaker-level embeddings (speaker_embeddings/)
The goal is to provide a turn-key dataset forforced alignment, prosody modelling, TTS, and speaker adaptation experiments without… See the full description on the dataset page: https://huggingface.co/datasets/srinathnr/TTS_DATASET.Somali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
Gusii-ASR-Data-Subset-470H
Gusii ASR Data Subset 470H
Gusii speech dataset for automatic speech recognition.
Kamba-ASR-Data-Subset-484H
Kamba ASR Data Subset 484H
Kamba speech dataset for automatic speech recognition.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.ASR-datasets-ptbr
📚 Datasets de Áudio em Português (PT-BR)
Este repositório reúne diversos corpora públicos de fala em português do Brasil, combinados em um único dataset para facilitar treinamentos e pesquisas em ASR (Automatic Speech Recognition).
O objetivo é fornecer um recurso amplo, padronizado e de fácil acesso para a comunidade.
📂 Datasets Integrados
A tabela abaixo lista todos os datasets incluídos, com suas informações:
Dataset
Config Name
TOTAL
train
test
validation… See the full description on the dataset page: https://huggingface.co/datasets/opedromartins/ASR-datasets-ptbr.Thinkspark-v2-270m-training-data
ThinkSpark-v2-350M — training data
Full-duplex floor-controller (Section 8) training corpus: playable audio + text,
paired for the Dataset Viewer, plus every scenario field (behaviour, language, domain,
gender, prosody, agent text) and Soniox character-level timestamps.
Dataset Viewer
Default split is parquet with a real Audio feature — a player renders inline next to
the text in the Hub UI:
column
type
description
audio
Audio
playable wav (already… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/Thinkspark-v2-270m-training-data.Emilia-Dataset-JA-Plus
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2024/08/28: Welcome to join Amphion's Discord channel to stay connected and engage with our community!
2024/08/27: The Emilia dataset is now publicly available! Discover the most extensive and diverse speech generation dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/Emilia-Dataset-JA-Plus.Audio-FLAN-Dataset
Audio-FLAN Dataset (Paper)
(the FULL audio files and jsonl files are still updating)
An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound.
1. Dataset Structure
The Audio-FLAN-Dataset has the following directory structure:
Audio-FLAN-Dataset/
├── audio_files/
│ ├── audio/
│ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/
│ │ └── 179_Audioset_for_Audio_Inpainting/
│ │ └── ...
│ ├── music/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak.
The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.VoiceIsolation-Benchmark-Dataset
Voice Isolation Benchmark Dataset
265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing.
265 recordings · 47 speakers · 3 scenarios · 65 scripts
Why this dataset exists
Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.MyMentorLLM-dataset
Dataset Card for MyMentorLLM
This dataset accompanies the MyMentorLLM paper, arXiv:2607.25667, which describes the simulation environment, generation procedure, experimental design and validation analyses. If you use this dataset, please cite the paper (see Citation Information).
Dataset Summary
MyMentorLLM is a multimodal voice- and text-based deliberate-practice environment used to generate 2,100 simulated Cognitive Behavioural Therapy (CBT) training sessions… See the full description on the dataset page: https://huggingface.co/datasets/RodolfoRizzi/MyMentorLLM-dataset.linto-dataset-audio-ar-tn
LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task
This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT
(linagora/linto-asr-ar-tn).
Dataset Summary
Dataset composition
Sources
Data Table
Data sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.khmer-speech-dataset
Khmer ASR Cultural Dataset
727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided.
Language: Khmer (khm).
Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata.
Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/phonsobon/khmer-speech-dataset.X-Voice-Dataset-Train
X-Voice Training Dataset
Overview
The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech models. It provides a robust foundation for cross-lingual phonetic and prosodic modeling.
Also the train set of X-Voice Model.
Core Statistics
Total Speech Duration: 420K hours
30 languages
European: bg (Bulgarian), cs (Czech), da (Danish), de (German), el (Greek), en (English), es (Spanish), et (Estonian), fi… See the full description on the dataset page: https://huggingface.co/datasets/XRXRX/X-Voice-Dataset-Train.
