Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fixie-ai /common_voice_17_0audio10M<n<100M20 likes159k downloads2y agoHugging Face02ropedia-ai /xperience-10mgated ⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Dataset Summary Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.3dvideo-classification1M<n<10M250 likes69k downloads6mo agoHugging Face03fixie-ai /covost2This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer. The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger. As such, not all the data is included: Only the validation and test subsets are available. From the XX_EN subsets, only fr, es, and zh-CN are included. audio1M<n<10M5 likes62k downloads2y agoHugging Face04AISHELL /AISHELL-3AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-3.audiotext-to-speech10K<n<100K16 likes28k downloads3y agoHugging Face05ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M124 likes21k downloads4mo agoHugging Face06shenyunhang /AISHELL-4 AISHELL-4 Identifier: SLR111 Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd Category: Speech License: CC BY-SA 4.0 Downloads (use a mirror closer to you): train_L.tar.gz [7.0G] &nbsp; ( Training set of large room, 8-channel microphone array speech ) &nbsp; Mirrors: [US] &nbsp; [EU] &nbsp; [CN] &nbsp; train_M.tar.gz [25G] &nbsp; ( Training set of medium room, 8-channel microphone array speech ) &nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-4.audio10K<n<100K2 likes12k downloads2y agoHugging Face07qyang1021 /AIR-Bench-Dataset AIR-Bench Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions. The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon). Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audioquestion-answeringn<1K8 likes10k downloads2y agoHugging Face08moonshine-ai /audio_samples_1kaudio0 likes9k downloads7mo agoHugging Face09AISHELL /AISHELL-4audio4 likes7k downloads3y agoHugging Face10ai4bharat /Kathbathgated Kathbath Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India Languages Bengali Gujarati Kannada Hindi Malayalam Marathi Odia Punjabi Sanskrit Tamil Telugu Urdu Licensing Information The IndicSUPERB dataset is released under this licensing scheme: We do not own any of the raw text used in creating this dataset. The text data… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Kathbath.audio100K<n<1M33 likes6.9k downloads2y agoHugging Face11malaysia-ai /malaysian-youtube Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs data at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube/data Notebooks at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube How to load the data efficiently? import pandas as pd import json from datasets import Audio from torch.utils.data import DataLoader, Dataset chunks = 30 sr = 16000 class Train(Dataset):… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-youtube.audio10K<n<100K5 likes6.5k downloads2y agoHugging Face12shenyunhang /AISHELL-3 AISHELL-3 Identifier: SLR93 Summary: Mandarin data, provided by Beijing Shell Shell Technology Co., Ltd. Category: Speech License: Apache License v.2.0 Downloads (use a mirror closer to you): data_aishell3.tgz [19G] &nbsp; (speech data and transcripts ) &nbsp; Mirrors: [US] &nbsp; [EU] &nbsp; [CN] &nbsp; About this resource:AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-3.audio10K<n<100K2 likes6.4k downloads2y agoHugging Face13ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M47 likes6.2k downloads2y agoHugging Face14ai4bharat /Rasagated Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies Overview We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.audiotext-to-speech1M<n<10M62 likes5.5k downloads4mo agoHugging Face15tarteel-ai /everyayah﷽ Dataset Card for Tarteel AI's EveryAyah Dataset Dataset Summary This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in Arabic. Dataset Structure Data Instances A typical data point comprises the audio file audio, and its transcription called text. The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.audioautomatic-speech-recognition100K<n<1M44 likes4.6k downloads23d agoHugging Face16akatz-ai /H3-Character-Swap-v1 H3 Character Swap v1 A reference-conditioned character-replacement dataset for MiniMax H3 Ref2VA LoRA training with Ostris AI Toolkit. It combines synthetic still-image edits with unchanged real-motion regularization videos. 134 examples: 94 character-swap edits and 40 preservation clips. Training has 76 edits + 32 clips; validation has 18 edits + 8 clips. Prepared resolution is 1344×768 at 24 fps. The companion 1,000-step LoRA are available separately. Task and… See the full description on the dataset page: https://huggingface.co/datasets/akatz-ai/H3-Character-Swap-v1.imagen<1K9 likes4.5k downloads15d agoHugging Face17hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes4.4k downloads4y agoHugging Face18ai4bharat /Shrutilipigated Shrutilipi Overview Shrutilipi is a labelled ASR corpus obtained by mining parallel audio and text pairs at the document scale from All India Radio news bulletins for 12 Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu. The corpus has over 6400 hours of data across all languages. This work is funded by Bhashini, MeitY and Nilekani Philanthropies Usage The datasets library… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Shrutilipi.audio1M<n<10M19 likes4.2k downloads2y agoHugging Face19EZMONYI /music-ai-human-test-audio Interpretable AI and Human Music Evaluation Archive Research audio and versioned experiment outputs for an English graduation thesis. The audio archive is incomplete. Completed experiments and verified partial audio publications must not be confused with whole-project delivery completion. No blanket license is assigned to this mixed-source archive. Completed experiments and thesis The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.audio1K<n<10K0 likes4.1k downloads27d agoHugging Face20fixie-ai /gigaspeechaudio10M<n<100M12 likes3.9k downloads2y agoHugging Face21shekar-ai /Neyshekar Neyshekar Neyshekar is an open, community-driven Persian speech dataset collected via a web-based crowdsourcing platform at https://ney.shekar.io. It is designed to support research and development in text-to-speech (TTS), automatic speech recognition (ASR), speech representation learning, and other downstream Persian speech applications. The recordings are provided by volunteer contributors, all of whom are native Persian speakers. Each release represents a… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/Neyshekar.audioautomatic-speech-recognition10K<n<100K2 likes2.1k downloads26d agoHugging Face22tarteel-ai /tloggatedTLOG is a dataset consisting of audio recitations of Quranic Ayahs and their corresponding Quranic texts. Features Each data point of this dataset consists of the following features: audio: audio recitation of a specific Ayah in the Quran for a clean data point: array: the audio signal in array form sample_rate: the sample rate of the audio signal path: the file name, which, for a clean data point, should correspond to the verse that’s being recited: The format is:… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/tlog.audio100K<n<1M21 likes2k downloads23d agoHugging Face23awakening-ai /ReactNet ResponseNet ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants. Paper If you use this dataset, please cite: ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ReactNet.audio1K<n<10K3 likes2k downloads1y agoHugging Face24laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M16 likes2k downloads11mo agoHugging Face25basis-ai /basis-conversations-1500gated Dataset Card for Basis Conversations 1500 Listen first: sample conversations Overview Basis Conversations 1500 is a multi-party, multilingual, full duplex conversational speech dataset. Each conversation includes up to 4 simultaneous speakers, each with channel-separated, 48 kHz audio. The median conversation lasts 33 minutes and 2,645 unique speakers are represented. Multi-party: a conversation seats between two and four people at a time; participants come… See the full description on the dataset page: https://huggingface.co/datasets/basis-ai/basis-conversations-1500.audioautomatic-speech-recognition10K<n<100K21 likes2k downloads17d agoHugging Face26besimple-ai /duplex-cue Duplex Cue Duplex Cue is an audio benchmark for evaluating how an ongoing speaker responds when another speaker contributes during the turn. It separates the listener's local intent from the ongoing speaker's observed behavior, allowing systems to be evaluated for both content uptake and floor management. The accompanying manuscript reports the complete study of 80 conversations and 300 canonical trials. This public repository preserves the manuscript and the full research code… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/duplex-cue.audioaudio-classificationn<1K3 likes1.8k downloads9d agoHugging Face27disco-eth /AIMEfrom datasets import load_dataset dataset = load_dataset('disco-eth/AIME') AIME: AI Music Evaluation Dataset The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo. The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset. The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset. The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.audio1K<n<10K9 likes1.7k downloads2y agoHugging Face28tarteel-ai /EA-DIaudio100K<n<1M8 likes1.7k downloads4y agoHugging Face29besimple-ai /voice-code-bench VoiceCodeBench VoiceCodeBench is a test-only benchmark for evaluating whether automatic speech recognition (ASR) systems preserve exact structured values in English workplace speech. Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition The benchmark targets cases where a transcript is software input: callback numbers, email addresses, command-line flags, file paths, URLs, account identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.audioautomatic-speech-recognitionn<1K14 likes1.7k downloads22h agoHugging Face30krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes1.7k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.