Team Ai
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cloud0day3 /alania-speech-captions-tr Alania Turkish Speech Style Captions English · Türkçe 835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses, volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in espnet/yodas3 (Turkish), so you fetch it from there. Instruction-following and voice-description TTS need… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-speech-captions-tr.tabulartext-to-speech100K<n<1M10 likes1.7k downloads9d agoHugging Face02surindersinghssj /gurbani-sehajpath-yt-captions-canonical Gurbani Sehajpath — Canonical-aligned ASR corpus Stage-1 + Stage-2 canonical pipeline output for sehaj-path (calm recitation of the Guru Granth Sahib). Built from publicly available audio recordings with aligned transcripts, chunked by caption timing and aligned against the canonical Guru Granth Sahib Ji text (SGGS). Columns Schema is auto-inferred from the parquet shards. Primary columns: audio — 16 kHz mono waveform final_text — canonical Gurmukhi transcription (post… See the full description on the dataset page: https://huggingface.co/datasets/surindersinghssj/gurbani-sehajpath-yt-captions-canonical.audioautomatic-speech-recognition10K<n<100K0 likes220 downloads6mo agoHugging Face03TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes142 downloads7mo agoHugging Face04TTS-AGI /majestrino-unified-detailed-captions-temporal Majestrino Unified Detailed Captions with Temporal Aspects Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects. Stats 4,128,665 samples 826 tar files (~1.1 GB each) ~878 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption with temporal aspects caption_type — always unified_detailed_caption_with_temporal_aspects transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioaudio-classification1M<n<10M0 likes100 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.