laion/laion-voice-profiles-annotated
Synthetic Voice-Profile Performances Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.
Synthetic Voice-Profile Performances
Authors: Christoph Schuhmann and LAION.
28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d speaker/style embedding, and 16 procedurally generated captions.
The default loader view is generated voice-profile takes and identity-repair rerenders (original + repair): 28,212,933 rows and 71,056 published view-hours. The repair rows are replacements, so adding 54,943 and 16,113 does not measure 71,056 independent final performances. The selected voice-converted profile takes form a different, overlapping derivative view; their 54,811 hours must not be added to the default view as new source performances. These descriptive names are labels only. The stored origin values, vc_sidon configuration and file paths remain unchanged for compatibility.
All three runs are complete. vc_sidon/ ships one clip per source take — the candidate the producer ranked first — so it has the same row count as origin=original minus the 736 takes whose winning clip the annotation pass dropped. It lives in its own directory and its own loader config (vc_sidon), so load_dataset(...) with the default config keeps returning exactly the 28,212,933 origin=* rows. See `vc_sidon/README.md`.
Licence and attribution
Released under CC-BY-4.0. If you use this data, credit Christoph Schuhmann and LAION, and the upstream sources below.
The 500 voice profiles were selected, not invented — drawn from a consolidated reference pool of 6,064 voices by a quality gate, a DNSMOS floor, stratified gender- and language-balanced quotas, and farthest-point sampling in a 10-dimensional VoiceNet summary space, so the 500 are maximally different from one another rather than the highest-scoring.
Only the emolia_* family traces back to real recorded human speech, and that upstream dataset is itself public under CC-BY-4.0. The other four families are model-reinterpreted new speakers whose source identifiers are deliberately opaque and are not reconstructible from this release. The audio here is in all cases model output, not an upstream recording. The consolidated reference pool the 500 were selected from is not public.
Every row retains a src_meta block carrying the complete original per-take generation metadata.
What a voice profile is
One reference speaker pushed through a large, fixed, named matrix of acting conditions:
origin=repair is a partial re-do of takes that failed an identity check in origin=original, not an independent sample. A repair row supersedes the original row with the same (voice, gid, audio_key).
Reconstructing either run
import pyarrow.dataset as ds
d = ds.dataset("index", partitioning="hive", format="parquet")
base_only = d.to_table(filter=ds.field("origin") == "original")
repaired_only = d.to_table(filter=ds.field("origin") == "repair")One rendering per utterance — base, with repair substituted where base failed the floor:
import pyarrow.dataset as ds
df = ds.dataset("index", partitioning="hive", format="parquet").to_table(
columns=["uid","voice","gid","audio_key","origin","caption_general"]).to_pandas()
df["_pref"] = (df["origin"] == "repair").astype(int)
one = (df.sort_values("_pref", ascending=False)
.drop_duplicates(["voice","gid","audio_key"], keep="first"))text, rank, empty, take, diff and tail shadow pandas DataFrame attributes — always use df["text"], never df.text.
Layout
README.md
code/ caption_render.py, caption2.py, mp3io.py, build_*.py
index/origin=original/part-*.parquet 2,000 files, 205 columns
index/origin=repair/part-*.parquet 500 files, 205 columns
scores/origin=original/*.parquet 83-column generation-run scores
manifests/shards.parquet source shard -> rows, bytes, sha256
data/origin=original/*.tar 2,000 WebDataset shards
data/origin=repair/chunk=NN/*.tar 29,956 shards across 6 chunk directories
stats/stats_profiles.csv 998 rows: per (voice, origin)
stats/stats_emotions.csv 80 rows: per (emotion, origin)
stats/stats_dimensions.csv 114 rows: per (VoiceNet dimension, origin)
stats/norm_stats_vprof.json population distributions for optional normalisation
vc_sidon/data/*.tar 2,000 WebDataset shards, winners only, 4,364 GB
vc_sidon/index/origin=vc_sidon/part-*.parquet one row per winner, 187 columns
vc_sidon/scores/origin=vc_sidon/part-*.parquet one row per VC candidate (all four), 149 columns
vc_sidon/manifests/ shards.parquet + sources.json
vc_sidon/README.md its own card -- read it before using that directoryTwo layout notes. data/origin=repair/ is split across chunk=00 … chunk=05 purely as a storage detail — the Hub caps any single directory at 10,000 files and that run has 29,956 shards. chunk carries no meaning; glob data/origin=repair/*/*.tar. Correspondingly, index/origin=repair/ ships as 500 consolidated parquets rather than one per shard, which is both a better read size and within the same cap. The shard column still names the tar each row came from, so the index-to-audio mapping is unchanged.
Where the audio is, and in what format
The audio is in the tars under `data/`, and nowhere else. No parquet in this repository has an audio column. Four members per utterance, sharing one uid stem:
The MP3s are two-channel files carrying duplicated mono. Read back from the frame headers of the shipped bytes in both partitions, not inferred. This is the single most useful fact about the audio in this repository, because it is exactly the shape that caused the half-speed bug described under Known issues below: a decoder that flattens stereo frames instead of taking one channel returns a signal exactly twice as long, at half speed, and every derived value computed from it is wrong while every internal consistency check still passes. soundfile, librosa, torchaudio and ffmpeg all do the right thing. A hand-rolled frame reader may not.
The MP3 bytes themselves were never affected and were never regenerated — only derived values were, and they have been recomputed.
Only in the tar, never in the parquet: words (word timestamps), src_meta, and the per-dimension softmax confidence voicenet[DIM]["conf"]. The JSON's top-level keys are uid, dataset, src_shard, dur_s, sr_src, lang, text, text_field, text_source, words, align_status, bursts, n_bursts, burst_placement, text_with_bursts, voicenet, genuineness_0_6, blend_0_10, emonet, quality, caption_general, caption_script, src_meta, extra.
The join — index row to audio
shard is the tar basename without the extension; uid is the member stem. Verified member for member on vprof_base-00000.tar and vprof_repaired-00000.tar.
import io, json, tarfile, glob
import numpy as np, pyarrow.dataset as ds
d = ds.dataset("index", partitioning="hive", format="parquet")
row = d.head(1).to_pylist()[0]
if row["origin"] == "original":
tar = f"data/origin=original/{row['shard']}.tar"
else: # repair is spread over chunk=00 .. chunk=05
tar = glob.glob(f"data/origin=repair/*/{row['shard']}.tar")[0]
with tarfile.open(tar) as tf:
stem = row["uid"] # the member stem, verbatim
audio = tf.extractfile(stem + ".mp3").read() # 48 kHz, 160 kbps, 2-channel
rec = json.loads(tf.extractfile(stem + ".json").read())
codes = np.load(io.BytesIO(tf.extractfile(stem + ".moss.npy").read())) # uint16 [T, 12]
vclap = np.load(io.BytesIO(tf.extractfile(stem + ".vclap.npy").read())) # float16 [768]
assert codes.shape == (row["moss_frames"], row["moss_n_vq"])
print(rec["words"][:3], row["caption_general"])uid contains dots — do not split on the first one. And `audio_key` is not unique across runs: join on uid, or on the triple (voice, gid, audio_key).
Column dictionary
205 columns in `index/`, 83 in `scores/`, 7 in `manifests/shards.parquet`. The complete generated reference — every column, its type and a one-line meaning, including all 40 emo_* and all 114 vn_* with the dimension each code stands for — is in [`COLUMNS.md`](COLUMNS.md). The groups below are the tour.
Identity
audio_key is not unique across runs — join on (voice, gid, audio_key), or use uid.
Audio and text
dur_s (seconds), lang (en/de), text (the generation prompt), text_source (source on every row), n_words, align_status, audio_bytes, moss_frames, moss_n_vq.
All parquet files across both partitions share one identical 205-column schema, so pyarrow.dataset discovery works over the whole index without schema promotion.
moss_frames == round(dur_s × 12.5) holds on 100.000% of rows (the tokenizer rounds; a floor rule matches only ~94.7% and is the wrong invariant).
Vocal bursts
n_bursts, burst_labels (list), burst_starts/burst_ends (lists, seconds), burst_placement (inline / general / none), text_with_bursts — the transcript with (burst:LABEL) inserted between the correct two words. null when placement is not inline.
Scores
emo_* ×40 (Empathic-Insight intensities, 0–7), qual_* ×4, genuineness_0_6, blend_0_10, top_emotion, top_emotion_value, vn_<CODE>_reg ×57 (unclamped regression estimate), vn_<CODE>_bucket ×57 (classification argmax; 0…6, BKGN 0…4, EXPL 0…2).
Model chain: laion/voiceclap-commercial (frozen, 768-d, audio padded/truncated to exactly 30 s at 16 kHz) → laion/voicenet-dimension-predictors-commercial.
VoiceNet reliability, stated plainly. These heads were trained on Gemini-3.5-flash perceptual estimates — not ground truth, not human ratings. Mean held-out MAE 0.744 scale-points, mean Pearson r 0.79, mean bucket accuracy 0.617 (chance ≈0.14 at k=7). Best:EXPLMAE 0.444,R_HEAD0.478,BKGN0.493. Worst:S_ASMR1.316,ARSH1.257,S_RANT1.238. `R_MIXD` has r = 0.39 — its labels collapsed and it is close to uninformative. Do not treat the weak axes as measurements. Levels are defined by prose, not a shared numeric anchor, so "high" on one dimension is not commensurate with "high" on another.
scores/ — generation-run scores (origin=original only)
An 83-column table produced during generation on correctly-decoded audio. Join on audio_key. Includes spk_sim, wer, genuineness, blend, reward, rank, is_holdout, in_train_pool, and matrix descriptors emotion, condition, dim, level, block, character, burst_class. There is no equivalent table for `origin=repair`; that run's generation metrics live only in src_meta inside the tars.
Captions
16 wordings per utterance as caption_<template>. caption_general is the pipeline's own output and is identical to caption_clausal — verified on 361,277 rows after the 2026-08-23 regeneration, which is what makes the two code paths (capfix.py for the column, caption_render.py for the templates) safe to keep separate. All 16 wordings share one emotion gate; see below.
They vary in clause order, verbosity, register, and burst handling, so a caption-conditioned model sees real paraphrase rather than reshuffling.
What the ordinal words are relative to — read before conditioning
The VoiceNet words are absolute; the emotion clause is corpus-relative. These are different and the difference is deliberate.
VoiceNet ordinal words (vn_<CODE>_bucket → "bright", "measured", "masculine") are absolute. The bucket is the argmax of an ordinal classification head whose levels were each defined by a paragraph of prose in absolute terms, and the renderer maps bucket -> word through a static table using no dataset statistics at all. "bright" means bright on the model's own scale, not "bright for this corpus". style_thr=3.0 and the explicit-content flag are fixed constants too. The caveat is upstream: those anchors are Gemini-3.5-flash estimates, so "absolute" means consistent with the model's training anchors, not calibrated to human ratings.
The emotion clause (reads as …) is relative to the wider corpus, as of 2026-08-23. It previously used an absolute emo_thr=1.0 across 40 Empathic-Insight heads that do not share a scale, which named Interest on 90.6 % of all captions and Bitterness on 0.1 % — it was reporting the scale of the head, not the emotion of the clip. An emotion is now named when it falls in the top 10 % for that emotion (U ≥ 0.90, at most 3, ranked by percentile), against code/capnorm.npz: a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning 8 datasets and every language. A clip that clears nothing says `no dominant emotion`. So "reads as sadness" means "unusually sad for this corpus", and Interest now appears on 5.3 % of rows. Full rationale and the gate constants: code/CAPTION_GATE.md.
Consequences worth knowing before you condition on these captions:
- The emotion clause is not comparable across corpora the way the VoiceNet words are. It is a percentile within the LAION-TTS corpus, so it is deliberately not re-normalised per dataset — doing that would force every dataset to the same emotional profile by construction.
- `no dominant emotion` is a real signal, not a missing value. It fires when a clip is unremarkable on all 40 heads. Dropping those rows biases a training set toward expressive speech.
stats/norm_stats_vprof.jsonstill ships per-dimension and per-emotion distributions computed from this release's own output. Use it if you want normalisation scoped to the voice profiles rather than to the whole corpus; do not mix the two scales in one comparison.
Regenerating or adding a wording
import pyarrow.dataset as ds, importlib.util
spec = importlib.util.spec_from_file_location("cr", "code/caption_render.py")
cr = importlib.util.module_from_spec(spec); spec.loader.exec_module(cr)
row = ds.dataset("index", partitioning="hive", format="parquet").head(1).to_pylist()[0]
cr.render(row, template="casting") # any of the 16
cr.render(row, template="clausal", polarity="v1_buggy") # the pre-fix wordingcaption_render.py renders from the flat numeric columns and nothing else — no audio, no GPU, no model, ~25 µs per caption. Add your own by putting a function in cr.TEMPLATES; it receives cr.features(row), which has already resolved every ladder, threshold and ranking.
Known issues and corrections
1. A half-speed decode bug was found and fully corrected before release
An earlier annotation pass decoded these files with a routine that flattened stereo frames. The audio is duplicated mono in a 2-channel MP3, so every sample appeared twice and the decoded signal was exactly 2× too long — half speed. Every GPU stage consumed it: the VoiceCLAP embedding (and therefore all 57 VoiceNet dimensions, genuineness, blend), the BUD-E heads (40 emo_*, 4 qual_*), burst detection, forced alignment, and the MOSS tokenizer.
The entire annotation stack was recomputed from the fixed decoder. Verification on the shipped data:
The MP3 audio was never affected and was not regenerated — the pipeline copies source bytes into the output tar unchanged. Only derived values were wrong, and they have been recomputed.
Why it hid: the internal check moss_frames == floor(dur_s × 12.5) passed on affected rows because both terms were doubled.
2. Caption polarity — corrected
The GEND and BKGN caption ladders ran backwards relative to the data.
- `GEND`:
vn_GEND_regcorrelates +0.842 withvn_R_CHST_reg(chest resonance, a masculine marker), −0.484 withvn_R_HEAD_reg, −0.352 withvn_BRGT_reg. HighGENDis masculine — and the model's own documentation agrees (bucket 0 hyper-feminine, top bucket hyper-masculine). The old ladder captioned the most masculine voices "strongly feminine". - `BKGN`: correlates +0.699 with
vn_RCQL_regand +0.29 withqual_background_quality, an independent head. HighBKGNis cleaner — the model's documentation agrees (bucket 0 background overwhelms the speech, top acoustically dead). The old ladder captioned clean studio-synthetic audio "very noisy background".
Only the prose inverted — every numeric column was always correct. The pre-fix wording remains reproducible via caption_general_v1_wording and polarity="v1_buggy".
A fourth, independent confirmation of the direction arrived after this text was first written: classifying all 500 voice profiles from the numeric vn_GEND against each profile's own design-spec card_gender agrees on 89.2 % of 379 decided voices. An inverted ladder would score about 11 %.
Superseded claim, kept visible rather than deleted. This section used to say the fix changes only GEND and BKGN terms with zero other differences. That was true of the polarity fix alone and is no longer true of the current captions: the 2026-08-23 regeneration also replaced the emotion clause (issue 6). Across the corpus the two changes together touched 83,146,345 rows for GEND, 79,455,045 for BKGN, and the emotion clause on nearly all of them. Delivery, timbre, speech, affect, style, recording quality, the explicit flag and burst handling remain untouched, and that is asserted per shard.
6. The emotion clause was regenerated on 2026-08-23 — re-read your captions
Every caption in this release was re-rendered. Two things changed and nothing else:
- The emotion clause now uses a corpus-percentile gate instead of an absolute threshold (see Captions above).
Interestfell from being named on 90.6 % of rows to 5.3 %; all 40 emotions now occur; 17.9 % of rows sayno dominant emotion. - `GEND`/`BKGN` polarity (issue 2) is applied to the rendered strings.
The previous string is preserved. caption_general_v1 holds the pre-regeneration caption_general verbatim on every row, so the change is auditable row-by-row and reversible. Note the distinction from the older caption_general_v1_wording, which is a re-render of the pre-fix GEND/BKGN wording over the corrected numbers — not the literal previous string.
If you cached captions before 2026-08-23, rebuild rather than patch. Anything derived from the old text — SFT/DPO prompt strings, retrieval indices, caption-conditioned training sets — carries the inverted GEND/BKGN wording and an Interest-saturated emotion clause. Do not attempt to fix it with find-and-replace: recompute from vn_GEND_bucket / vn_BKGN_bucket and the emotion columns, which is what code/capfix.py does, and which is idempotent.
Determinism note: the ECDF table is kept in float64. An earlier float32 cast collapsed genuinely different percentiles onto one value (emo_Relief 0.99780922730 and emo_Contentment 0.99780920848 both became 0.99780923128), which made tied emotions order nondeterministically. Genuine ties now break by descending percentile, then ascending emotion name — reproducible from the caption alone.
3. There is no Parakeet ASR for these runs
If you want raw ASR with word timestamps: it does not exist here. What exists is MMS_FA forced alignment of the generation prompt text onto the audio — text_source is source on every row. The wer column in scores/ is Whisper large-v3-turbo, computed at generation time. nvidia/parakeet-tdt-0.6b-v3 is used elsewhere in the wider corpus but was never run over the voice profiles. About 1.4% of rows have a non-ok align_status (CTC length failures).
4. Synthetic data with a deliberate stratified bias
Every utterance is model output, not recorded human speech. The 842-condition matrix is a stratified design, not a natural distribution: emotions, voice qualities and edge cases are massively over-represented relative to natural speech. Do not compute population statistics from this corpus. The three vprof_* runs are near-duplicates of the same 500 voices saying the same lines — concatenating them without weighting over-represents this material.
5. The MOSS codes are a re-analysis, not the model's original codes
The audio was generated with MOSS — the model emitted codes and decoded them — but only the decoded MP3 was written. Re-encoding gives codes for the same audio, not the codes the model actually emitted.
Related
- Base model: `laion/moss-tts-local-transformer-4.55b-voice-acting-v2`
- LoRA adapters: `laion/moss-voice-profile-loras-500`
- Reference audios: `TTS-AGI/moss-voice-profile-references`
- Codec: `OpenMOSS-Team/MOSS-Audio-Tokenizer-v2`
Citation
@misc{schuhmann_laion_voice_profiles_2026,
title = {LAION Voice Profiles — Annotated},
author = {Schuhmann, Christoph and {LAION}},
year = {2026},
url = {https://huggingface.co/datasets/laion/laion-voice-profiles-annotated}
}