datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-acting-cutscene-prompts
Cut-Scene Voice-Acting Prompts
Continuously-generated, character-consistent two-scene "CUT TO:" voice-performance
prompts (text only, no audio) for training and evaluating expressive TTS / voice-acting
models. Each prompt describes a single speaker across two sharply contrasting emotional
moments separated by a CUT TO: transition, in a voice-acting stage-direction format
(spoken lines in "quotes", performance notes in (parentheses)).
Total prompts: 4,057,000
Languages: English… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-cutscene-prompts.laion-voice-profiles-annotated
Synthetic Voice-Profile Performances
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from
500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting
conditions. Every
utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads,
vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a
768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.sorrel-sft-voiceVoiceAssistant-400K-SLAM-Omni
VoiceAssistant-400K (Modified)
This dataset is prepared for the reproduction of SLAM-Omni.
This is a single-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/VoiceAssistant-400K-SLAM-Omni.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.wav2vec2_common_voice_accents_3laion-voice-profiles-dpo-cfg
LAION Voice Profiles — contrastive DPO pairs (CFG + phase 2)
Authors: Christoph Schuhmann and LAION.
842,935 preference pairs in four families, built from the same 500 synthetic voice profiles as
laion/laion-voice-profiles-sft
and laion/laion-voice-profiles-dpo.
These are the two pair families that the sister DPO set does not contain: they were built later,
for two measured defects of the models trained on it, and they are the complete remainder of the
project's preference… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo-cfg.Custom_common_voice_dataset_using_RVC
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
Custom common_voice_v11 corpus with a custom voice was was created using RVC(Retrieval-Based Voice Conversion)
The model underwent 200 epochs of training, utilizing a total of 1 hour of audio clips. The data was scraped from Youtube.
The audio in the custom generated dataset is of a YouTuber named
Ajay Pandey
Description
license: cc0-1.0
language:
- hi… See the full description on the dataset page: https://huggingface.co/datasets/Aniket-Tathe-08/Custom_common_voice_dataset_using_RVC.arknights_voices_zh
ZH Voice-Text Dataset for Arknights Waifus
This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12431 records, 25.9 hours in total. Average duration is approximately 7.49s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.laion-voice-profiles-dpo
LAION Voice Profiles — TTS preference pairs (DPO)
Authors: Christoph Schuhmann and LAION.
3,451,531 preference pairs in three families, built from the same 500 synthetic voice profiles
as laion/laion-voice-profiles-sft.
Every pair shares one prompt; chosen and rejected differ only in the way the family names.
config
pairs
teaches
how rejected is made
emotion
1,064,594
hit the right tone for this line
a take of the same sentence, same voice, rendered at the wrong… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-dpo.moss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2)
Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices.
Each prompt is a fixed, optimized champion performance direction (instruction) paired
with a Gemma-generated topic text (text) — together, one performance to render. For every
prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64
within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.accented_common_voicemoss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.Emilia-YODAS-Voice-Conversion
Emilia-YODAS-Voice-Conversion
We sample https://huggingface.co/datasets/amphion/Emilia-Dataset YODAS set for voice conversion.
Filter transcriptions based on character repetitiveness and word ngrams.
Filter speaker similarity using https://huggingface.co/nvidia/speakerverification_en_titanet_large during speaker permutation.
Convert audio to speech tokens using https://huggingface.co/neuphonic/neucodec
We also upload the full permutation as zip files.
Speech Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Emilia-YODAS-Voice-Conversion.APAC-Egocentric-Residential-Voiceover
APAC Egocentric Residential (with Voiceover)
Ten narrated first-person recordings of household chores, each shipping the original capture with spoken voiceover, a burned-in caption render, WebVTT captions, and an ASS annotation track.
This is the only release in the HumynLabs egocentric collection that carries audio narration — the wearer describes each action as they perform it, and the captions align that speech to the video.
Preview: 45 s from the cooking sample, captioned… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/APAC-Egocentric-Residential-Voiceover.common_voice_17_0
Common Voice Corpus 17.0
Mirror for mozilla-foundation/common_voice_17_0, easy to download and extract instead audio in parquet files.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_17_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
arknights_voices_jp
JP Voice-Text Dataset for Arknights Waifus
This is the JP voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
10905 records, 26.3 hours in total. Average duration is approximately 8.7s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_427_vigil_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_jp.azurlane_voices_jp
JP Voice-Text Dataset for Azur Lane Waifus
This is the JP voice-text dataset for azur lane playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
30160 records, 75.8 hours in total. Average duration is approximately 9.05s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/azurlane_voices_jp.common_voice_22_0
Common Voice Corpus 22.0
Originally from https://huggingface.co/datasets/fsicoli/common_voice_22_0, we mirror using multiple zip files also trimmed the silents.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_22_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/common_voice_22_0.emolia-voicenet-gemini-annotations
Emolia VoiceNet Gemini Annotations
468,180 dimension-level annotations over 236,613 Emolia speech clips,
each scored 0-6 (0-2 for the content-safety dimension) on one of 57 perceptual
voice / speech dimensions - arousal, valence, brightness, resonance placement, speaking
styles, genuineness, recording quality, and more - by Gemini 3.5 Flash (non-thinking,
temperature 0). This repository ships the annotations, audio provenance, per-dimension
statistics, and the full scoring… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-voicenet-gemini-annotations.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.prompt-voice-v1.5
Dataset Overview
This dataset contains nearly 2.35M English speech instruction to text answer samples, using the combination of:
Intel/orca_dpo_pairs
routellm/gpt4_dataset
nomic-ai/gpt4all-j-prompt-generations
microsoft/orca-math-word-problems-200k
allenai/WildChat-1M
Open-Orca/oo-gpt4-200k
Magpie-Align/Magpie-Pro-300K-Filtered
qiaojin/PubMedQA
Undi95/Capybara-ShareGPT
HannahRoseKirk/prism-alignment
BAAI/Infinity-Instruct
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/prompt-voice-v1.5.laion-voice-profiles-sft
LAION Voice Profiles — TTS supervised fine-tuning set
Authors: Christoph Schuhmann and LAION.
1,200,531 instruction-tuning samples for reference-conditioned TTS, drawn from 500 synthetic
voice profiles: for each of the 842 acting conditions of each voice, the best 3 of its 48
candidate takes, each paired with a different clip of the same voice from another group as
the reference, plus the exact conditioning prompt, MOSS audio codes for both target and reference,
word-level… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-sft.VoiceTrace-BenchVoiceTrace-Bench
VoiceTrace is a benchmark and unified framework for who-said-what speech retrieval: given a natural-language query about a speaker's identity or what they said, retrieve the matching audio document. Unlike conventional speaker verification or diarization benchmarks, VoiceTrace evaluates retrieval jointly over who is speaking and what is being said, across both single-speaker and multi-speaker conversational recordings.
This repository hosts the VoiceTrace-Bench… See the full description on the dataset page: https://huggingface.co/datasets/cara-ai/VoiceTrace-Bench.ko-voicephishing-binary-classificationVoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series
multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.
