datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_voice_17_0xperience-10m
⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified.
Interactive Intelligence from Human Xperience
Xperience-10M
Dataset Summary
Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.covost2This is a partial copy of CoVoST2 dataset.
The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer.
The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger.
As such, not all the data is included: Only the validation and test subsets are available.
From the XX_EN subsets, only fr, es, and zh-CN are included.
AISHELL-3AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-3.IndicVoices
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
Updates
[23 December 2025] We now have 11,200 hours of transcribed data! 🎉
Overview
INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.AISHELL-4
AISHELL-4
Identifier: SLR111
Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd
Category: Speech
License: CC BY-SA 4.0
Downloads (use a mirror closer to you):
train_L.tar.gz [7.0G] ( Training set of large room, 8-channel microphone array speech
) Mirrors:
[US]
[EU]
[CN]
train_M.tar.gz [25G] ( Training set of medium room, 8-channel microphone array speech
) … See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-4.AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audio_samples_1kAISHELL-4Kathbath
Kathbath
Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India
Languages
Bengali
Gujarati
Kannada
Hindi
Malayalam
Marathi
Odia
Punjabi
Sanskrit
Tamil
Telugu
Urdu
Licensing Information
The IndicSUPERB dataset is released under this licensing scheme:
We do not own any of the raw text used in creating this dataset.
The text data… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Kathbath.malaysian-youtube
Malaysian Youtube
Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours.
URLs data at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube/data
Notebooks at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube
How to load the data efficiently?
import pandas as pd
import json
from datasets import Audio
from torch.utils.data import DataLoader, Dataset
chunks = 30
sr = 16000
class Train(Dataset):… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-youtube.AISHELL-3
AISHELL-3
Identifier: SLR93
Summary: Mandarin data, provided by Beijing Shell Shell Technology Co., Ltd.
Category: Speech
License: Apache License v.2.0
Downloads (use a mirror closer to you):
data_aishell3.tgz [19G] (speech data and transcripts
) Mirrors:
[US]
[EU]
[CN]
About this resource:AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus
published by Beijing Shell Shell Technology Co.,Ltd. It can be… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-3.indicvoices_r
IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS
Dataset Summary
IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.Rasa
Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages
Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies
Overview
We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering
a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.everyayah﷽
Dataset Card for Tarteel AI's EveryAyah Dataset
Dataset Summary
This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters.
Supported Tasks and Leaderboards
[Needs More Information]
Languages
The audio is in Arabic.
Dataset Structure
Data Instances
A typical data point comprises the audio file audio, and its transcription called text.
The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.H3-Character-Swap-v1
H3 Character Swap v1
A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos.
134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately.
Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.genshin-voice-v3.3-mandarin
Dataset Card for Genshin Voice
Dataset Description
Dataset Summary
The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game.
Languages
The text in the dataset is in Mandarin.
Dataset Creation
Source Data
Initial Data Collection and Normalization
The data was obtained by unpacking the Genshin Impact game.
Who are the source language producers?
The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.Shrutilipi
Shrutilipi
Overview
Shrutilipi is a labelled ASR corpus obtained by mining parallel audio and text pairs at the document scale from All India Radio news bulletins for 12 Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu. The corpus has over 6400 hours of data across all languages.
This work is funded by Bhashini, MeitY and Nilekani Philanthropies
Usage
The datasets library… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Shrutilipi.music-ai-human-test-audio
Interpretable AI and Human Music Evaluation Archive
Research audio and versioned experiment outputs for an English graduation thesis.
The audio archive is incomplete. Completed experiments and verified partial
audio publications must not be confused with whole-project delivery completion.
No blanket license is assigned to this mixed-source archive.
Completed experiments and thesis
The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.gigaspeechNeyshekar
Neyshekar
Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at https://ney.shekar.io. It is designed to support research and development in text-to-speech (TTS), automatic speech recognition (ASR), speech representation learning, and other downstream Persian speech applications.
The recordings are provided by volunteer contributors, all of whom are native Persian speakers. Each release represents a… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/Neyshekar.tlogTLOG is a dataset consisting of audio recitations of Quranic Ayahs and their corresponding Quranic texts.
Features
Each data point of this dataset consists of the following features:
audio: audio recitation of a specific Ayah in the Quran
for a clean data point:
array: the audio signal in array form
sample_rate: the sample rate of the audio signal
path: the file name, which, for a clean data point, should correspond to the verse that’s being recited:
The format is:… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/tlog.ReactNet
ResponseNet
ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants.
Paper
If you use this dataset, please cite:
ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ReactNet.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
basis-conversations-1500
Dataset Card for Basis Conversations 1500
Listen first: sample conversations
Overview
Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented.
Multi-party: a conversation seats between two and four people at a time; participants come… See the full description on the dataset page: https://huggingface.co/datasets/basis-ai/basis-conversations-1500.duplex-cue
Duplex Cue
Duplex Cue is an audio benchmark for evaluating how an ongoing speaker responds when another speaker contributes during the turn. It separates the listener's local intent from the ongoing speaker's observed behavior, allowing systems to be evaluated for both content uptake and floor management.
The accompanying manuscript reports the complete study of 80 conversations and
300 canonical trials. This public repository preserves the manuscript and the
full research code… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/duplex-cue.AIMEfrom datasets import load_dataset
dataset = load_dataset('disco-eth/AIME')
AIME: AI Music Evaluation Dataset
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset.
The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset.
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.EA-DIvoice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.VoiceAgentBench
VoiceAgentBench
This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978).
VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.
