datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.6k-diverse-reference-voices
6k Diverse Reference Voices
6,064 permissively licensed reference voices for casting expressive voice-acting generations.
All voices in this collection are permissively usable: they were either synthetically created or
extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0.
Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0),
re-published under LAION with clarified metadata documentation. If you use this dataset,
please attribute… See the full description on the dataset page: https://huggingface.co/datasets/laion/6k-diverse-reference-voices.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.moss-reference-voices-consolidated
MOSS reference voices — consolidated (6,064 voices)
6,064 reference voices for casting MOSS-VA-v2 voice-acting generations. Every voice was auto-annotated
by Gemini (name, tagline, language, accent, age/gender read, register, timbre, distinctive features,
emotional range, casting suggestions for 4 genres, free-text tags, search text) and scored on 99 measured
dimensions: 57 VoiceNet voice-quality axes (timbre/prosody/register/speaking-style, e.g. brightness,
roughness, warmth… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-reference-voices-consolidated.
