datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alania-synthetic-speech-tr
Alania Turkish Synthetic Speech
English · Türkçe
3,411 hours of 48 kHz Turkish speech in 2,037,461 clips, spoken by 2,752 designed voices plus
one-off voices, each clip paired with the written sentence, its spoken form and a plain-English description of the
voice. Every clip is AI-generated: no real person's voice is in this dataset.
We made it while building Alania-2, the Turkish text-to-speech model behind
speech.patientdesk.ai. Openly licensed Turkish speech for TTS is… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-synthetic-speech-tr.alania-speech-captions-tr
Alania Turkish Speech Style Captions
English · Türkçe
835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses,
volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the
measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in
espnet/yodas3 (Turkish), so you fetch it from there.
Instruction-following and voice-description TTS need… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/alania-speech-captions-tr.antalia-voice-corpus
Antalia Turkish single-speaker scripted speech
5.008 hours / 1,073 clips of studio-quality Turkish read speech from one professional voice
actor, script-aligned, segmented, and quality-gated. This is the corpus the
Antalia 1 voice was fine-tuned on.
Clean, consented, single-speaker Turkish speech at this quality is scarce — which is the main
reason this release exists. Development of the model is discontinued; the data is published
as-is so it stays useful.
Model:… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-voice-corpus.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.
