Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AquaV /genshin-voices-separated21 likes235k downloads2y agoHugging Face02fixie-ai /common_voice_17_0audio10M<n<100M18 likes210k downloads2y agoHugging Face03VoiceOfML /VOMEBOOK 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/VOMEBOOK/discussions 提出。 此仓库存储马列之声电子书:https://huggingface.co/datasets/VoiceOfML/VOMEBOOK/tree/main 。 请使用:https://voiceofml-search.hf.space/VOMEBOOK 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/VOMEBOOK )。 可使用:https://voiceofml-search.hf.space/VOMEBOOK?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/VOMEBOOK?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/VOMEBOOK.document10K<n<100K8 likes63k downloads2d agoHugging Face04fsicoli /common_voice_15_0 Dataset Card for Common Voice Corpus 15.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 15. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_15_0.automatic-speech-recognition100B<n<1T6 likes39k downloads3y agoHugging Face05fsicoli /common_voice_22_0 Dataset Card for Common Voice Corpus 22.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_22_0.automatic-speech-recognition100B<n<1T20 likes28k downloads1y agoHugging Face06VoiceHub /voicehub-arena-seed-tts-eval VoiceHub Arena — native TTS evaluations Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. Interactive demo · Source code Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.audiotext-to-speech0 likes25k downloads16d agoHugging Face07simon3000 /genshin-voice Genshin Voice Genshin Voice is a dataset of voice lines from the popular game Genshin Impact. Hugging Face 🤗 Genshin-Voice ModelScope Genshin-Voice Per-speaker downloads are grouped by language and ZIP size. Browse every archive in the ZIP index. Last update at 2026-08-13 654252 wavs 7291 without speaker (1%) 52693 without transcription (8%) 1088 without inGameFilename (0%) Dataset Details Dataset Description The dataset contains voice lines… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/genshin-voice.audioaudio-classification100K<n<1M281 likes17k downloads1mo agoHugging Face08laion /voice-acting-cutscene-prompts Cut-Scene Voice-Acting Prompts Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting models. Each prompt describes a single speaker across two sharply contrasting emotional moments separated by a CUT TO: transition, in a voice-acting stage-direction format (spoken lines in "quotes", performance notes in (parentheses)). Total prompts: 4,057,000 Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.tabulartext-generation1M<n<10M2 likes17k downloads25d agoHugging Face09fsicoli /common_voice_17_0 Dataset Card for Common Voice Corpus 17.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 17. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_17_0.automatic-speech-recognition100B<n<1T19 likes15k downloads2y agoHugging Face10fsicoli /common_voice_16_0 Dataset Card for Common Voice Corpus 16.0 This dataset is an unofficial version of the Mozilla Common Voice Corpus 16. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/. Languages Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese (China), Chinese (Hong Kong), Chinese (Taiwan), Chuvash, Czech… See the full description on the dataset page: https://huggingface.co/datasets/fsicoli/common_voice_16_0.automatic-speech-recognition100B<n<1T4 likes14k downloads3y agoHugging Face11ebook2audiobook /E2A-Voices0 likes13k downloads24d agoHugging Face12VoiceOfML /MLMRL-Hub 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/discussions 提出。 此仓库存储马列毛主义与革命左翼仓储中心和图书馆的未重复资料(仅经一次md5检测,压缩包内容未去重。):https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub/tree/main 。 请使用https://voiceofml-search.hf.space/MLMRL-Hub 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Hub )。 可使用:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Hub?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Hub.audio10K<n<100K0 likes13k downloads1d agoHugging Face13mteb /common_voice_21_00 likes12k downloads1y agoHugging Face14VoiceOfML /Teachers 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Teachers/discussions 提出。 此仓库存储导师著作:https://huggingface.co/datasets/VoiceOfML/Teachers/tree/main 。 请使用:https://voiceofml-search.hf.space/Teachers 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Teachers )。 可使用:https://voiceofml-search.hf.space/Teachers?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Teachers?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without large files - just their… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Teachers.document1K<n<10K2 likes10k downloads23d agoHugging Face15legacy-datasets /common_voiceCommon Voice is Mozilla's initiative to help teach machines how real people speak. The dataset currently consists of 7,335 validated hours of speech in 60 languages, but we’re always adding more voices and languages.automatic-speech-recognition100K<n<1M148 likes9.1k downloads2y agoHugging Face16VoiceOfML /SovMaterials 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/SovMaterials/discussions 提出。 此仓库存储苏联资料:https://huggingface.co/datasets/VoiceOfML/SovMaterials/tree/main 。 请使用:https://voiceofml-search.hf.space/SovMaterials 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/SovMaterials )。 可使用:https://voiceofml-search.hf.space/SovMaterials?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/SovMaterials?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone without… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/SovMaterials.document1K<n<10K5 likes8.2k downloads23d agoHugging Face17hlt-lab /voicebench License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in your research, please cite the following paper: @article{chen2024voicebench, title={VoiceBench: Benchmarking LLM-Based Voice Assistants}, author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou}, journal={arXiv preprint arXiv:2410.17196}, year={2024} } audio10K<n<100K16 likes7.2k downloads1y agoHugging Face18ducthinh8283 /voicetts7 likes7.1k downloads1mo agoHugging Face19simon3000 /zenless-voice Zenless Voice Zenless Voice is a dataset of voice lines from the popular game Zenless Zone Zero. Hugging Face 🤗 Zenless-Voice ModelScope Zenless-Voice Per-speaker downloads are grouped by language and WAV count. Browse every archive in the ZIP index. Last update at 2026-09-17, game version 3.2.0 406720 wavs 78785 without speaker (19%) 123429 without transcription (30%) 83509 without inGameFilename (21%) Speaker archives contain 327,935 WAVs in 4,322 ZIPs. The 78,785 rows… See the full description on the dataset page: https://huggingface.co/datasets/simon3000/zenless-voice.audioaudio-classification100K<n<1M10 likes7k downloads18d agoHugging Face20hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes6.4k downloads4y agoHugging Face21VoiceOfML /MLMRL-Library 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/discussions 提出。 此仓库存储重要书库备份:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/tree/main 。 请使用:https://voiceofml-search.hf.space/MLMRL-Library 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Library )。 可使用:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Library.document1 likes6.4k downloads2d agoHugging Face22VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes5.2k downloads3mo agoHugging Face23ayf3 /numberblocks-one-voice-datasetaudio1K<n<10K7 likes5k downloads4d agoHugging Face24laion /laion-voice-profiles-annotated Synthetic Voice-Profile Performances Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.tabulartext-to-speech10M<n<100M0 likes4.9k downloads11d agoHugging Face25gpt-omni /VoiceAssistant-400Kaudio100K<n<1M100 likes4.9k downloads2y agoHugging Face26AquaV /fallout-4-voicesaudio10K<n<100K3 likes4.5k downloads2y agoHugging Face27VoiceOfML /Japanese-Materials 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/Japanese-Materials/discussions 提出。 此仓库存储日共资料:https://huggingface.co/datasets/VoiceOfML/Japanese-Materials/tree/main 。 请使用:https://voiceofml-search.hf.space/Japanese-Materials 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/Japanese-Materials )。 可使用:https://voiceofml-search.hf.space/Japanese-Materials?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/Japanese-Materials?wide=1 )。… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/Japanese-Materials.audion<1K0 likes4.5k downloads23d agoHugging Face28VoiceOfML /GPCREducation 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/GPCREducation/discussions 提出。 此仓库存储文化大革命教材:https://huggingface.co/datasets/VoiceOfML/GPCREducation/tree/main 。 请使用:https://voiceofml-search.hf.space/GPCREducation 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/GPCREducation )。 可使用:https://voiceofml-search.hf.space/GPCREducation?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/GPCREducation?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/GPCREducation.document1K<n<10K1 likes4.4k downloads4mo agoHugging Face29krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes4.1k downloads8mo agoHugging Face30UncovAI /Real_Voiceaudio100K<n<1M1 likes3.8k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.