text-to-speech
text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.Tamazight-Speech-to-Arabic-Text
Tamazight-Arabic Speech Recognition Dataset
This is the Tamazight-NLP organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset. This dataset contains ~15.5 hours of Tamazight (Tachelhit dialect) speech paired with Arabic transcriptions, designed for automatic speech recognition (ASR) and speech-to-text translation tasks.
Dataset Details
Total Examples: 20,344 audio segments
Training Set: 18,309 examples (~8.9GB)
Test Set: 2,035 examples (~992MB)… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Tamazight-Speech-to-Arabic-Text.700h-tr-turkish-text-to-speechdarija_speech_to_textspeech_to_text_yixing_dialectDataset-Text-To-Speech-Indonesia
🎵 Dataset Audio Bahasa Indonesia
Dataset audio berkualitas tinggi untuk Text-to-Speech (TTS) bahasa Indonesia.
Dibuat oleh : Muhammad Arief, S.Kom.Universitas Muhammadiyah SorongTeknik Informatika 2020
📊 Spesifikasi Teknis
Parameter
Nilai
Satuan
Total Durasi
16.38
jam
Jumlah Segmen
4531
file
Durasi Rata-rata
13.01
detik
Sample Rate KHz
22
kHz
Sample Rate Hz
22000
Hz
Bit Depth
PCM_16
PCM
Format
wav
Lossless
🔄 Urutan Pengolahan… See the full description on the dataset page: https://huggingface.co/datasets/X-lord/Dataset-Text-To-Speech-Indonesia.
