laion/6k-diverse-reference-voices
6k Diverse Reference Voices 6,064 permissively licensed reference voices for casting expressive voice-acting generations. All voices in this collection are permissively usable: they were either synthetically created or extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0. Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0), re-published under LAION with clarified metadata documentation. If you use this dataset, please attribute… See the full description on the dataset page: https://huggingface.co/datasets/laion/6k-diverse-reference-voices.
6k Diverse Reference Voices
6,064 permissively licensed reference voices for casting expressive voice-acting generations.
All voices in this collection are permissively usable: they were either synthetically created or extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0.
Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0), re-published under LAION with clarified metadata documentation. If you use this dataset, please attribute both LAION and the original TTS-AGI source.
Every voice was auto-annotated by Gemini (name, tagline, language, accent, age/gender read, register, timbre, distinctive features, emotional range, casting suggestions for 4 genres, free-text tags, search text) and scored on 99 measured dimensions: 57 VoiceNet voice-quality axes, 40 Empathic-Insight emotion axes, plus genuineness and burst-blend. Each voice ships with three audio variants and per-variant DNSMOS.
Live search UI: https://projects.laion.ai/moss-reference-voice-search/ — code + reproducer under `search_tool/` in this repo.
Composition
Counts by source (= per-voice source field, also the cid prefix). Verified from metadata.parquet:
Language (Gemini-identified, 4,727/6,064 non-null): English 4,250 · German 473 · null 1,337 · rest 4. Gender read: Male 4,262 · Female 1,675 · Androgynous 109 · Non-human 8 · edge cases 10.
The three audio variants — and why
Each voice has three parallel renders of the same demo clip:
Dataset-level mean DNSMOS-OVRL is tied (orig 3.343, sidon 3.346, cbx 3.344 — see annotations/dnsmos_stats.json), but per-voice ~62% win with a processed variant (orig wins 2,317 = 38.2%, sidon 1,940 = 32.0%, cbx 1,807 = 29.8%). Use the precomputed `best_version` field (argmax DNSMOS per voice) instead of defaulting to one variant dataset-wide.
File layout
data/voices-0000.tar … voices-0011.tar # WebDataset shards, ~505-506 voices each, ~2 GB total
metadata.parquet # flat index, one row per voice (6064 x 30, see below)
annotations/
dims.npy # (6064, 99) float32 — 99-dim scores on ORIG audio
dims_enh.npy # (6064, 99) float32 — 99-dim scores on SIDON audio
dim_catalog.json # 99-dim schema: [{i, code, name, group, desc}]
dnsmos.json # {cid: {orig, sidon, cbx}} DNSMOS-OVRL per variant
dnsmos_stats.json # means + per-variant win counts / win_pct
search_tool/ # FastAPI search server + pipeline + demo page
README.md / LICENSE # this file / CC-BY-4.0WebDataset shards (data/*.tar)
Members for one voice are contiguous:
<cid>.orig.mp3 # original demo clip
<cid>.sidon.mp3 # SIDON-denoised + loudness-normalized
<cid>.cbx.mp3 # Chatterbox self-conversion of sidon
<cid>.json # full per-voice record (see below)The <cid>.json record = full voice entry (name, tagline, gender, age, language, accent, register, timbreprofile, distinctivefeatures, emotionalrange, casting {classicfantasy, scifi, mysteryhorror, contemporary}, tags, search_text, legacy scores, source) plus:
{
"dnsmos": {"orig": 3.44, "sidon": 3.40, "cbx": 3.29},
"best_version": "orig",
"dims_raw": [99 floats, order = annotations/dim_catalog.json, scored on orig],
"dims_enh": [99 floats, same order, scored on sidon]
}import webdataset as wds
ds = wds.WebDataset("hf://datasets/LAION/6k-diverse-reference-voices/data/voices-{0000..0011}.tar").decode()
for sample in ds:
cid = sample["__key__"]
rec = sample["json"] # dict with dnsmos, best_version, dims_raw, dims_enh
print(cid, rec["name"], rec["best_version"])metadata.parquet — flat index (1 row / voice)
Columns (30 total):
- Identity / casting:
cid, name, gender, age, language, accent, tagline, tags (list<string>), source, shard - Quality per variant:
dnsmos_orig, dnsmos_sidon, dnsmos_cbx(float, DNSMOS-OVRL 0–5, higher = better),best_version(orig|sidon|cbx= argmax DNSMOS) - Key 99-dim values with both scorings (
dim_*= on orig,dim_*_enh= on sidon):dim_GEND (+_enh)perceived gender (higher = more masculine),dim_AGEV (+_enh)perceived age,dim_GENU (+_enh)genuineness (sounds like real human recording),dim_BLEND (+_enh)vocal-burst blend quality,dim_BKGN (+_enh)background-noise level,dim_VALN (+_enh)/dim_AROU (+_enh)emotional valence / arousal,dim_WARM (+_enh)vocal warmth.
import pandas as pd
df = pd.read_parquet("metadata.parquet")
loud_masculine = df[(df.dim_GEND > 4) & (df.dnsmos_orig > 3.4)]Row order of metadata.parquet == row order of annotations/dims.npy / dims_enh.npy.
99-dim scores (annotations/)
dim_catalog.json: list of 99{i, code, name, group, desc},i= column index into the.npyfiles. Groups:emonet40 (indices 0–39, higher = stronger emotion),voicenet57 (indices 40–96, timbre/prosody/register/style),quality2 (index 97GENUgenuineness, 98BLENDvocal-burst blend).dims.npy: float32(6064, 99), scored on orig audio.dims_enh.npy: same shape, scored on sidon audio. No NaNs. Observed range approx −3.9 … 12.2 (raw regressor outputs, not clipped to 0–6).- Per-tar
dims_raw== correspondingdims.npyrow;dims_enh==dims_enh.npyrow. dnsmos.json:{cid: {"orig": float, "sidon": float, "cbx": float}}for all 6,064 voices.
License
CC-BY-4.0. You may share and adapt with attribution to LAION and the original TTS-AGI/moss-reference-voices-consolidated source.
Search
See `search_tool/` for the FastAPI server (BM25 / sentence-embedding / VoiceCLAP text→audio similarity, optional AND-filters over any of the 99 dims), the pipeline scripts (SIDON, Chatterbox self-conversion, DNSMOS, 99-dim scoring, VoiceCLAP, assembly), and the live-demo reproducer.
