Team Ai
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /experiment-speaker-embeddingaudion<1K0 likes942 downloads2y agoHugging Face02humair-experiments /genshin-voice-english-fXLaudio100K<n<1M0 likes283 downloads5mo agoHugging Face03humair-experiments /tts-pretrain-clones-3m TTS Pretrain Clones (3M) 2,967,779 clone utterances across 2971 English speakers. Sample rate: 44.1 kHz, WAV in Parquet Generated by echo-tts synthesizing English text on speaker latents derived from Qwen3-TTS VoiceDesign base speakers. Per speaker: 10 voice-clone latents × 100 texts. The first utterance of each speaker (row 0) is published separately in the companion refs set. Coverage: speakers 1-60 + 61 (partial, 749 rows) + 91-3000. Thirty speakers (61's tail + 62-90) are… See the full description on the dataset page: https://huggingface.co/datasets/humair-experiments/tts-pretrain-clones-3m.audiotext-to-speech1M<n<10M0 likes198 downloads5mo agoHugging Face04inaam1995 /commonvoice17_su_experiments Common Voice 17 -- Single / Long Utterance experiment dataset Built from fixie-ai/common_voice_17_0 (English); the original CV splits are preserved and each is bucketed into single-utterance (1 word) and long-utterance (>= 3 words). Splits: dev_single, dev_long, test_single, test_without_single, train_single, train_long. audio1M<n<10M0 likes187 downloads4mo agoHugging Face05humair-experiments /genshin-voice-english-fXaudio1K<n<10K0 likes19 downloads5mo agoHugging Face06Mathani-Ayat /qaloon-reciter-experiments Experimental only — Waleed and Trabulsi Private research inventory uploaded at the project owner's request. Not part of the approved training dataset. Both sources retain authorization_status: unauthorized; an experimental upload is not source redistribution permission. Waleed: 569 individually labelled WAVs, durations read from WAV headers; Fātiḥah uses a separate basmalah track and fused ending. Source labels have not been independently checked. Some source identity entries… See the full description on the dataset page: https://huggingface.co/datasets/Mathani-Ayat/qaloon-reciter-experiments.audioautomatic-speech-recognition1K<n<10K0 likes17 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.