datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audiofolder_two_configs_in_metadataaudiofolder_single_config_in_metadataaudiofolder_no_configs_in_metadataAudioSet
Dataset Card for AudioSet
Dataset Summary
AudioSet is a dataset of 10-second clips from YouTube, annotated into one or more sound categories, following the AudioSet ontology.
Supported Tasks and Leaderboards
audio-classification: Classify audio clips into categories. The leaderboard is available here
Languages
The class labels in the dataset are in English.
Dataset Structure
Data Instances
Example… See the full description on the dataset page: https://huggingface.co/datasets/agkphysics/AudioSet.open-asr-leaderboard
ESB Test Sets: Parquet & Sorted
This dataset takes the open-asr-leaderboard/datasets-test-only data and sorts each split by audio length.
The format is also changed, from custom loading script (un-safe remote code) to parquet (safe).
Broadly speaking, this dataset was generated with the following code-snippet:
from datasets import load_dataset, get_dataset_config_names
DATASET = "open-asr-leaderboard/datasets-test-only" # dataset to load from
HUB_DATASET_ID =… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard.unwebtv-multilingual-audio-archive
UN WebTV Multilingual Audio Archive
This public dataset is a reproducibility-oriented archive of multilingual audio tracks collected from publicly accessible institutional media pages.
Files are grouped by source and published as source-level archives. The accompanying manifests preserve source URLs, language labels, extraction status, and verification metadata. The archive is intended for research and engineering evaluation; downstream users must respect the terms, licenses… See the full description on the dataset page: https://huggingface.co/datasets/hysi-lab/unwebtv-multilingual-audio-archive.audiofolder_two_configs_in_metadataNatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.audiofolder_two_configs_in_metadata_with_defaultdummy-audio-samplesLAION-Audio-300MKling-Audio-Eval
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
🌐 Website | 📖 arXiv
📋 Dataset Structure
The dataset structure is as follows:
Kling-Audio-Eval
├── Folder (first-level label)
│ ├── Folder (second-level label)
│ │ ├── video
│ │ │ └── *.mp4
│ │ ├── audio
│ │ │ └── *.wav
│ │ └── caption.csv # Header: video, audio, audio_tag, video_caption… See the full description on the dataset page: https://huggingface.co/datasets/klingfoley/Kling-Audio-Eval.stg-paired-audioaudio-course-imagesopen-asr-leaderboard-resultsaudio_samples_1kpsg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.avspeech-visual-audio
AVSpeech Video + Audio
This repository is a media-bearing reconstruction of the public AVSpeech
annotations. Each row represents an already-trimmed segment and keeps the
original source-video timing and target-face-center metadata.
Dataset structure
clip_id: identifier derived as
{youtube_id}_{start_sec:.3f}_{end_sec:.3f}.
avspeech_metadata: JSON containing youtube_id, start_sec, end_sec,
x_center, and y_center from the AVSpeech annotation.
video: video-only… See the full description on the dataset page: https://huggingface.co/datasets/ProgramComputer/avspeech-visual-audio.AudioSet
AudioSet
Google's AudioSet with the audio: 1,780,876 of its 2,084,320 labelled 10 s YouTube segments as 48 kHz FLAC, each with a
record of where its audio came from, measured quality, duplicate and eval-overlap flags, and a reason for every segment
that could not be found.
clips with audio
hours
classes
format
size
1,780,876 of 2,084,320 (85.4%)
4,904
527
48 kHz, 24-bit FLAC
2.25 TiB
The 48 kHz release in audio48k/ is the only audio in this repository. The… See the full description on the dataset page: https://huggingface.co/datasets/Muno459/AudioSet.big_bench_audio
Artificial Analysis Big Bench Audio
Dataset Summary
Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input.
The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022):
Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions
Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.AudiobookRu
AudiobookRu
Russian-language audiobook audio (chapter-level), scraped from the web. Each row is one chapter: embedded MP3 bytes + book/chapter metadata (~96k available audiobooks, 892,699 chapters).
Chapter-level fields: id (audiofile id), book_id (parent book), chapter_order, chapter_title, n_chapters, connected_book_id (linked text edition when present). main_actor_name is the narrator.
Quranic-Translation-Audio-Data
Overview
Quranic Translation Audio Data is a highly curated, standardized, and streaming-optimized multilingual audio dataset containing the complete recitation of translation audios and commentaries of the Holy Quran across 51 different translation directories.
Every audio track has been meticulously converted from heavy .mp3 source files into the modern, high-fidelity Opus (.opus) format at a streaming-optimized bitrate of 32kbps. Alongside… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Translation-Audio-Data.AudioSet_unbalanced_videoaudiosnippetsAudio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.Quranic-Word-By-Word-Audio-Data
🌟 Overview
Quran Word-By-Word Audio Dataset contains two complete word-by-word recitation datasets of the Holy Quran, optimized for edge delivery, mobile streaming, and machine learning pipelines:
Muallim (Teacher Style) — optimized for slow, educational, and repeat-friendly listening.
Mujawwad (Tajweed Style) — optimized for natural rhythmic recitation with full tajweed flow.
Originally averaging between 2.0 GB to 2.3 GB each in raw format, the… See the full description on the dataset page: https://huggingface.co/datasets/zaibihassan/Quranic-Word-By-Word-Audio-Data.persian-asr-audio-text-2.69M-chizzled
🗂️ persian-asr-audio-text-2.69M-chizzled
English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission
🌟 At a glance | معرفی سریع
English
فارسی
🎯 Purpose
Phase A-scale audio/text dataset.
پیکرهٔ بزرگ جفتهای صوت و متنِ پالایششده برای آموزش در مقیاس فاز A.
🧩 Role
Persian text and linguistic asset
مصنوع متنی و زبانی فارسی
📦 Snapshot
417 files; approximately 236.86 GB
417 فایل؛ حدود 236.86 GB
🧱 Packaging
414 Parquet files and 0… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-audio-text-2.69M-chizzled.Multitask-National-Speech-Corpus-v1-extendAudioCapsAudioSet-Strong
