datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdpval_preference_rubricstext-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.preference-asr-bench
Preference-ASR
Preference-ASR tests whether an ASR system
can follow natural-language instructions about how a transcript should be
written. Should numbers appear as words or digits? Should disfluencies stay in
the transcript? How should names be spelled and cased?
Conventional ASR benchmarks assume a single answer style, and their normalizers
often erase the differences those questions are meant to test. Each example in
Preference-ASR instead pairs an audio clip with a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/preference-asr-bench.tts-human-preferences-large
TTS Human Preferences (Large)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the large (2,700-row) subset. See also: small (1,000 rows), medium (2,000 rows).
Dataset Summary
Metric
Value
Total rows
2,700
Annotations per row
15
Total annotations
40,500
Unique… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-large.synthetic_audio_paired_preferences
Synthetic Audio Paired Preferences
A synthetic spoken-dialogue DPO (Direct Preference Optimization) dataset with 8482 examples,
designed for preference alignment of textless speech language models.
Dataset Summary
Each example corresponds to a single assistant turn in a two-speaker conversation.
For every turn the dataset provides:
The spoken prompt (the last user utterance, as audio + text)
A chosen response: the best model-generated spoken reply (scored by an LLM… See the full description on the dataset page: https://huggingface.co/datasets/vprak17/synthetic_audio_paired_preferences.tts-human-preferences-small
TTS Human Preferences (Small)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the small (1,000-row) subset. Larger versions will follow.
Dataset Summary
Metric
Value
Total rows
1,000
Annotations per row
15
Total annotations
15,000
Unique prompts
1,000
Audio format… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-small.tts-human-preferences-medium
TTS Human Preferences (Medium)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the medium (2,000-row) subset. See also: small (1,000 rows). Larger versions will follow.
Dataset Summary
Metric
Value
Total rows
2,000
Annotations per row
15
Total annotations
30,000… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-medium.merge_preference_data
