datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audiofolder_two_configs_in_metadataaudiofolder_single_config_in_metadataaudiofolder_no_configs_in_metadataAudioSet
Dataset Card for AudioSet
Dataset Summary
AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology.
Supported Tasks and Leaderboards
audio-classification: Classify audio clips into categories. The leaderboard is available here
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.open-asr-leaderboard
ESB Test Sets: Parquet & Sorted
This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length.
The format is also changed, from custom loading script (un-safe remote code) to parquet (safe).
Broadly speaking, this dataset was generated with the following code-snippet:
from datasets import load_dataset, get_dataset_config_names
DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from
HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.unwebtv-multilingual-audio-archive
UN WebTV Multilingual Audio Archive
This public dataset is a reproducibility-oriented archive of multilingual audio tracks collected from publicly accessible institutional media pages.
Files are grouped by source and published as source-level archives. The accompanying manifests preserve source URLs, language labels, extraction status, and verification metadata. The archive is intended for research and engineering evaluation; downstream users must respect the terms, licenses… See the full description on the dataset page: https://huggingface.co/datasets/hysi-lab/unwebtv-multilingual-audio-archive.NatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audiofolder_two_configs_in_metadataaudiofolder_two_configs_in_metadata_with_defaultdummy-audio-samplesopen-asr-leaderboard-resultsKling-Audio-Eval
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
🌐 Website | 📖 arXiv
📋 Dataset Structure
The dataset structure is as follows:
Kling-Audio-Eval
├── Folder (first-level label)
│ ├── Folder (second-level label)
│ │ ├── video
│ │ │ └── *.mp4
│ │ ├── audio
│ │ │ └── *.wav
│ │ └── caption.csv # Header: video, audio, audio_tag, video_caption… See the full description on the dataset page: https://huggingface.co/datasets/klingfoley/Kling-Audio-Eval.stg-paired-audiopsg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.audio-course-imagesaudio_samples_1kLAION-Audio-300Mavspeech-visual-audio
AVSpeech Video + Audio
This repository is a media-bearing reconstruction of the public AVSpeech
annotations. Each row represents an already-trimmed segment and keeps the
original source-video timing and target-face-center metadata.
Dataset structure
clip_id: identifier derived as
{youtube_id}_{start_sec:.3f}_{end_sec:.3f}.
avspeech_metadata: JSON containing youtube_id, start_sec, end_sec,
x_center, and y_center from the AVSpeech annotation.
video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.AudioSet_unbalanced_videoAudioSet
AudioSet
Google's AudioSet with the audio: 1,780,876 of its 2,084,320 labelled 10 s YouTube segments as 48 kHz FLAC, each with a
record of where its audio came from, measured quality, duplicate and eval-overlap flags, and a reason for every segment
that could not be found.
clips with audio
hours
classes
format
size
1,780,876 of 2,084,320 (85.4%)
4,904
527
48 kHz, 24-bit FLAC
2.25 TiB
The 48 kHz release in audio48k/ is the only audio in this repository. The… See the full description on the dataset page: https://huggingface.co/datasets/Muno459/AudioSet.AudiobookRu2
AudiobookRu2
Continuation of Muncy/AudiobookRu (that repo hit the HF storage quota). Chapter-level Russian audiobooks: embedded MP3 + book/chapter metadata. This repo contains the chapters that did not fit into part 1.
Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.AudioCapsQuranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.AudioMCQ-StrongAC-GeminiCoT
[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT
This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly.
Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis.
🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.AudiobookRu
AudiobookRu
Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata (~96k available audiobooks, 892,699 chapters).
Chapter-level fields: id (audiofile id), book_id (parent book), chapter_order, chapter_title, n_chapters, connected_book_id (linked text edition when present). main_actor_name is the narrator.
AudioSet-Stronglibrispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.myanmartts_generated_millions_audios
MyanmarTTS Generated Millions Audios
Status: generating in progress — more parquets arriving continuously.
Millions of Burmese sentences synthesized to audio using
MyanmarTTS — a from-scratch
31M-parameter flow-matching TTS model (CC0 1.0, pip install myanmartts).
Source Text
Derived from freococo/myanmar_spoken_corpus,
which merges Burmese text from FineWeb2, DCAI, DCAD200, and other web corpora.
Format
All files are Parquet, each holding ~5,000… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmartts_generated_millions_audios.
