datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
experiment-speaker-embeddinggenshin-voice-english-fXLtts-pretrain-clones-3m
TTS Pretrain Clones (3M)
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers. Per speaker:
10 voice-clone latents × 100 texts. The first utterance of each
speaker (row 0) is published separately in the companion refs set.
Coverage: speakers 1-60 + 61 (partial, 749 rows) + 91-3000. Thirty
speakers (61's tail + 62-90) are… See the full description on the dataset page: https://huggingface.co/datasets/humair-experiments/tts-pretrain-clones-3m.commonvoice17_su_experiments
Common Voice 17 -- Single / Long Utterance experiment dataset
Built from fixie-ai/common_voice_17_0 (English); the original CV splits are preserved and each is bucketed into single-utterance (1 word) and long-utterance (>= 3 words).
Splits: dev_single, dev_long, test_single, test_without_single, train_single, train_long.
genshin-voice-english-fXqaloon-reciter-experiments
Experimental only — Waleed and Trabulsi
Private research inventory uploaded at the project owner's request. Not part
of the approved training dataset. Both sources retain authorization_status: unauthorized; an experimental upload is not source redistribution permission.
Waleed: 569 individually labelled WAVs, durations read from WAV headers;
Fātiḥah uses a separate basmalah track and fused ending. Source labels have
not been independently checked. Some source identity entries… See the full description on the dataset page: https://huggingface.co/datasets/Mathani-Ayat/qaloon-reciter-experiments.
