datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.Quranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.AudiobookRu
AudiobookRu
Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata (~96k available audiobooks, 892,699 chapters).
Chapter-level fields: id (audiofile id), book_id (parent book), chapter_order, chapter_title, n_chapters, connected_book_id (linked text edition when present). main_actor_name is the narrator.
librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.AudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.Latin-Audio
Dataset Summary
Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training.
Alignment and curation: Kaiyuan Zhao
Language: Latin (Classical)
Uses
This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audio-v2This dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It wa released on April 20th, 2025.
You can find the full list of sources in this dataset under the dataset's sources.txt.
Paper: https://arxiv.org/abs/2307.08720
If you use our datasets, the following quote is preferable:
@misc{marmor2023ivritai,
title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development},
author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.t2a-audios-v1
t2a-audios-v1
The original text2asmr corpus (previously aoxo/audios): 48 kHz stereo ASMR audio with word-level
alignments, used for the v1 generator (Chatterbox speech LoRA, Stable Audio Open trigger LoRA) and as
the source for the reconstructed trigger ontology.
Superseded for ontology work by aoxo/t2a-mommy and
aoxo/t2a-daddy, which are larger, creator-attributed
and split by voice.
path
what
<id>.m4a
source audio, 48 kHz
<id>.json
word-level alignment + silence… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-audios-v1.lex-fridman-podcast-transcript-audio
Dataset Card for "lexFridmanPodcast-transcript-audio"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Lex Fridman Podcast. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Lex Fridman Podcast.
Data Fields
The dataset is composed by:
id: Id of the youtube… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/lex-fridman-podcast-transcript-audio.wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
police-scanner-audio
Police Scanner Audio Dataset
A comprehensive collection of police and emergency services radio communications from multiple US cities, captured from publicly available scanner feeds.
Dataset Overview
This dataset contains 103,660 audio recordings totaling 357GB of police scanner audio from 6 different cities across the United States. The recordings span multiple months of continuous monitoring and represent real-world emergency services communications.
Scanner… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/police-scanner-audio.CORAA-NURC-SP-Audio-Corpus
NURC-SP Corpus
NURC-SP Corpus CORAA ASR is a publicly available dataset for Automatic Speech Recognition (ASR) in the Brazilian Portuguese language containing 239.68 hours of audios (239.30 when filtered) and their respective transcriptions (170k+ segmented audios).
The audios were either validated by annotators or transcribed for the first time aiming at the ASR task.
Dataset Structure
The dataset is stored as Parquet files, with the audio embedded in each row… See the full description on the dataset page: https://huggingface.co/datasets/nilc-nlp/CORAA-NURC-SP-Audio-Corpus.Audio-FLAN-Dataset
Audio-FLAN Dataset (Paper)
(the FULL audio files and jsonl files are still updating)
An Instruction-Tuning Dataset for Unified Audio Understanding and Generation Across Speech, Music, and Sound.
1. Dataset Structure
The Audio-FLAN-Dataset has the following directory structure:
Audio-FLAN-Dataset/
├── audio_files/
│ ├── audio/
│ │ └── 177_TAU_Urban_Acoustic_Scenes_2022/
│ │ └── 179_Audioset_for_Audio_Inpainting/
│ │ └── ...
│ ├── music/
│ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Audio-FLAN-Dataset.dhivehi-audios-82-spk
Dhivehi Synthetic Voice and Speech Augmentation Dataset
This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-audios-82-spk.audio-v2-transcripts
Overview
This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license.
It was released on May 18th, 2025.
You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt.
All files were transcribed using the process.py pipeline, performing:
Frame-level VAD
Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.YouTube-Commons-nl-audio
YouTube Commons NL Audio
This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions,
all under a CC BY 4.0 license.
It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB.
Source
The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.seamless-interaction-audio-only
Seamless Interaction — audio only
Audio-only repackaging of Meta's
Seamless Interaction
dataset (improvised and naturalistic subsets), CC-BY-NC 4.0.
Changes from the original: video and motion (npz) data removed; audio
resampled from 48 kHz float to 24 kHz 16-bit PCM, stored as FLAC.
All label fields are kept as published:
vad, transcript: the per-file metadata/*.jsonl records.
annotation_1P_IS, annotation_1P_R, annotation_3P_IS, annotation_3P_R,
annotation_3P_V: the raw… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/seamless-interaction-audio-only.linto-dataset-audio-ar-tn
LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task
This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT
(linagora/linto-asr-ar-tn).
Dataset Summary
Dataset composition
Sources
Data Table
Data sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.infore2_audiobooks
unofficial mirror of InfoRe Technology public dataset №2
official announcement: https://www.facebook.com/groups/j2team.community/permalink/1010834009248719/
415h, 315k samples, vietnamese audiobooks of chinese wǔxiá 武俠 & xiānxiá 仙俠
bộ dữ liệu bóc ra từ YouTube đọc truyện võ hiệp & tiên hiệp, áp dụng kĩ thuật đối chiếu văn bản để dán nhãn tự động
official download:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/infore2_audiobooks.linto-dataset-audio-ar-tn-augmented
LinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT task
This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT linagora/linto-asr-ar-tn.
Dataset Summary
Dataset composition
Sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian Augmented is a dataset that builds on LinTO… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn-augmented.riigikogu-audio-stenograms-2018-2025
Riigikogu Stenograms 2018-2025
This dataset contains stenograms (approximate transcripts) of the sessions of the Estonian Parliament Riigikogu, spanning a time period from the very end of 2017 to May 2025, together with the corresponding audio.
The transcripts are not verbatim (word-by-word) transcripts but are edited for readability and grammatical correctness. Sentence start and end times (w.r.t. to the corresponding audio file) are provided.
The transcripts are converted into… See the full description on the dataset page: https://huggingface.co/datasets/TalTechNLP/riigikogu-audio-stenograms-2018-2025.radiotalk-us-audio-grok-clean
radiotalk-us-audio-grok-clean
Clean TTS audio for the v3 radiotalk transcripts: one row per transmission,
24 kHz mono PCM_16 WAV. Covers all 49,975 rendered scenarios of
twangodev/radiotalk-us-transcripts-grok-4.20-50k
— 527,701 utterances in uniform 1,250-row shards.
Synthesis: xAI Grok TTS API, 26 preset voices. Each scenario's speakers get
a deterministic voice assignment (seeded by scenario id) and a fixed
per-speaker speaking rate in 1.0–1.3×. 9 scenarios were dropped for… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-grok-clean.radiotalk-us-audio-tada-noisy
RadioTalk US Audio (Noisy)
VHF AM aviation channel-degraded variants of synthesized US air-traffic-control speech. One row per (clean turn × variant), embedded 8 kHz mono PCM_16 WAV.
This is the noisy variant of twangodev/radiotalk-us-audio-tada-clean — same transcripts and voices, passed through a probabilistic channel-simulation pipeline calibrated to the ATCO2 corpus SNR distribution (mean ~8 dB, range -5 to +30 dB) and shaped to ITU-R M.1084 / DO-186B aero voice passband… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-tada-noisy.ivirits-audio-v2-30s
ivrit.ai audio-v2 — 2–30 s segments
ivrit-ai/audio-v2 (>20k hours of Hebrew
audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR
fine-tuning.
How it was built
VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions
longer than 30 s are split at the quietest sufficiently-long pause inside the window,
so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped.
Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.radiotalk-us-audio-grok-noisy
radiotalk-us-audio-grok-noisy
VHF-AM channel-degraded counterpart to
twangodev/radiotalk-us-audio-grok-clean:
3 independently-degraded variants per clean utterance (bandpass, noise,
fading, heterodyne, PTT clicks, codec artifacts — the same radiotalk radio
pipeline behind the higgs/tada noisy sets). 1,583,103 rows covering all rendered scenarios of
twangodev/radiotalk-us-transcripts-grok-4.20-50k.
Difficulty (Grok STT)
On a 10,000-utterance sample, Grok STT scores… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/radiotalk-us-audio-grok-noisy.
