Team Ai
Datasetpublic

laion/laion-voice-profiles-annotated

Synthetic Voice-Profile Performances Authors: Christoph Schuhmann and LAION. 28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-voice-profiles-annotated.

sourceHugging Facecc-by-4.0updated 15d agoView on Hugging Face
0likes4.3kdownloads
Dataset Card

Synthetic Voice-Profile Performances

Authors: Christoph Schuhmann and LAION.

28,212,933 utterances / 71,056 hours of synthetic English and German voice-acting speech from 500 distinct voice profiles, each driven through the same fixed matrix of 842 named acting conditions. Every utterance carries 40 emotion intensities, 57 perceptual voice dimensions, 4 audio-quality heads, vocal-burst detections with timings, word-level forced alignment, MOSS audio codec tokens, a 768-d speaker/style embedding, and 16 procedurally generated captions.

Human-readable viewStored origin / loader IDWhat the rows representUtterancesPublished view-hoursVoices
Generated voice-profile takesorigin=original (vprof_base)Candidate takes from the initial synthesis run.20,079,38154,943500
Identity-repair rerendersorigin=repair (vprof_repaired)Replacement takes for failed identity checks; they supersede matching originals.8,133,55216,113498
Selected voice-converted profile takes (VC + restoration)vc_sidon (vprof_vc)One selected Chatterbox-VC candidate per source take, subsequently restored with SIDON.20,078,64554,811500

The default loader view is generated voice-profile takes and identity-repair rerenders (original + repair): 28,212,933 rows and 71,056 published view-hours. The repair rows are replacements, so adding 54,943 and 16,113 does not measure 71,056 independent final performances. The selected voice-converted profile takes form a different, overlapping derivative view; their 54,811 hours must not be added to the default view as new source performances. These descriptive names are labels only. The stored origin values, vc_sidon configuration and file paths remain unchanged for compatibility.

All three runs are complete. vc_sidon/ ships one clip per source take — the candidate the producer ranked first — so it has the same row count as origin=original minus the 736 takes whose winning clip the annotation pass dropped. It lives in its own directory and its own loader config (vc_sidon), so load_dataset(...) with the default config keeps returning exactly the 28,212,933 origin=* rows. See `vc_sidon/README.md`.


Licence and attribution

Released under CC-BY-4.0. If you use this data, credit Christoph Schuhmann and LAION, and the upstream sources below.

The 500 voice profiles were selected, not invented — drawn from a consolidated reference pool of 6,064 voices by a quality gate, a DNSMOS floor, stratified gender- and language-balanced quotas, and farthest-point sampling in a 10-dimensional VoiceNet summary space, so the 500 are maximally different from one another rather than the highest-scoring.

voice prefixvoicesupstreamlicencenature
emolia_c*219`laion/Emolia`cc-by-4.0real recorded multilingual emotional speech, speaker-clustered
mediathek_*124German public-broadcast clusters (source repo not public)—reinterpreted into new synthetic German speakers
k<n>_age<n>_bg<n>115clusters of generated audio—synthetic character voices
refvoice_*27clustered real + AI snippets—reinterpreted into new synthetic speakers
anime_*15anime-speech clusters—reinterpreted, synthetic, English delivery

Only the emolia_* family traces back to real recorded human speech, and that upstream dataset is itself public under CC-BY-4.0. The other four families are model-reinterpreted new speakers whose source identifiers are deliberately opaque and are not reconstructible from this release. The audio here is in all cases model output, not an upstream recording. The consolidated reference pool the 500 were selected from is not public.

Every row retains a src_meta block carrying the complete original per-take generation metadata.


What a voice profile is

One reference speaker pushed through a large, fixed, named matrix of acting conditions:

blockcodegroupsvaries
emotionE32040 emotions × 4 stage directions (A/B/C/D) × 2 languages
voicenetV45657 VoiceNet dimensions × 4 levels × 2 languages
edgeX2814 edge cases × 2 languages
characterC24character voices (e.g. "gravelly orc warlord")
burst_isolatedB10a single isolated vocal burst, no surrounding speech
sportsS2sports-commentary register
explicitP2safety-gated; legitimately absent for some voices

origin=repair is a partial re-do of takes that failed an identity check in origin=original, not an independent sample. A repair row supersedes the original row with the same (voice, gid, audio_key).

Reconstructing either run

python
import pyarrow.dataset as ds
d = ds.dataset("index", partitioning="hive", format="parquet")
base_only     = d.to_table(filter=ds.field("origin") == "original")
repaired_only = d.to_table(filter=ds.field("origin") == "repair")

One rendering per utterance — base, with repair substituted where base failed the floor:

python
import pyarrow.dataset as ds
df = ds.dataset("index", partitioning="hive", format="parquet").to_table(
        columns=["uid","voice","gid","audio_key","origin","caption_general"]).to_pandas()
df["_pref"] = (df["origin"] == "repair").astype(int)
one = (df.sort_values("_pref", ascending=False)
         .drop_duplicates(["voice","gid","audio_key"], keep="first"))

text, rank, empty, take, diff and tail shadow pandas DataFrame attributes — always use df["text"], never df.text.


Layout

README.md
code/          caption_render.py, caption2.py, mp3io.py, build_*.py
index/origin=original/part-*.parquet     2,000 files, 205 columns
index/origin=repair/part-*.parquet         500 files, 205 columns
scores/origin=original/*.parquet         83-column generation-run scores
manifests/shards.parquet                 source shard -> rows, bytes, sha256
data/origin=original/*.tar               2,000 WebDataset shards
data/origin=repair/chunk=NN/*.tar        29,956 shards across 6 chunk directories
stats/stats_profiles.csv                 998 rows: per (voice, origin)
stats/stats_emotions.csv                 80 rows: per (emotion, origin)
stats/stats_dimensions.csv               114 rows: per (VoiceNet dimension, origin)
stats/norm_stats_vprof.json              population distributions for optional normalisation
vc_sidon/data/*.tar                        2,000 WebDataset shards, winners only, 4,364 GB
vc_sidon/index/origin=vc_sidon/part-*.parquet   one row per winner, 187 columns
vc_sidon/scores/origin=vc_sidon/part-*.parquet  one row per VC candidate (all four), 149 columns
vc_sidon/manifests/                        shards.parquet + sources.json
vc_sidon/README.md                         its own card -- read it before using that directory

Two layout notes. data/origin=repair/ is split across chunk=00 … chunk=05 purely as a storage detail — the Hub caps any single directory at 10,000 files and that run has 29,956 shards. chunk carries no meaning; glob data/origin=repair/*/*.tar. Correspondingly, index/origin=repair/ ships as 500 consolidated parquets rather than one per shard, which is both a better read size and within the same cap. The shard column still names the tar each row came from, so the index-to-audio mapping is unchanged.

Where the audio is, and in what format

The audio is in the tars under `data/`, and nowhere else. No parquet in this repository has an audio column. Four members per utterance, sharing one uid stem:

memberformatdetail
<uid>.mp3MPEG-1 Layer III, 48 kHz, 160 kbps CBR, 2-channelthe generated take, source bytes copied through unchanged
<uid>.jsonJSONthe full nested record, including src_meta and words
<uid>.moss.npynumpy uint16, shape [T, 12]MOSS-Audio-Tokenizer-v2 codes. 12 codebooks x 1024, 12.5 fps, one frame = 12 tokens = 80 ms
<uid>.vclap.npynumpy float16, shape [768]the VoiceCLAP speaker/style embedding

The MP3s are two-channel files carrying duplicated mono. Read back from the frame headers of the shipped bytes in both partitions, not inferred. This is the single most useful fact about the audio in this repository, because it is exactly the shape that caused the half-speed bug described under Known issues below: a decoder that flattens stereo frames instead of taking one channel returns a signal exactly twice as long, at half speed, and every derived value computed from it is wrong while every internal consistency check still passes. soundfile, librosa, torchaudio and ffmpeg all do the right thing. A hand-rolled frame reader may not.

The MP3 bytes themselves were never affected and were never regenerated — only derived values were, and they have been recomputed.

Only in the tar, never in the parquet: words (word timestamps), src_meta, and the per-dimension softmax confidence voicenet[DIM]["conf"]. The JSON's top-level keys are uid, dataset, src_shard, dur_s, sr_src, lang, text, text_field, text_source, words, align_status, bursts, n_bursts, burst_placement, text_with_bursts, voicenet, genuineness_0_6, blend_0_10, emonet, quality, caption_general, caption_script, src_meta, extra.

The join — index row to audio

shard is the tar basename without the extension; uid is the member stem. Verified member for member on vprof_base-00000.tar and vprof_repaired-00000.tar.

python
import io, json, tarfile, glob
import numpy as np, pyarrow.dataset as ds

d   = ds.dataset("index", partitioning="hive", format="parquet")
row = d.head(1).to_pylist()[0]

if row["origin"] == "original":
    tar = f"data/origin=original/{row['shard']}.tar"
else:                                  # repair is spread over chunk=00 .. chunk=05
    tar = glob.glob(f"data/origin=repair/*/{row['shard']}.tar")[0]

with tarfile.open(tar) as tf:
    stem  = row["uid"]                                   # the member stem, verbatim
    audio = tf.extractfile(stem + ".mp3").read()          # 48 kHz, 160 kbps, 2-channel
    rec   = json.loads(tf.extractfile(stem + ".json").read())
    codes = np.load(io.BytesIO(tf.extractfile(stem + ".moss.npy").read()))   # uint16 [T, 12]
    vclap = np.load(io.BytesIO(tf.extractfile(stem + ".vclap.npy").read()))  # float16 [768]

assert codes.shape == (row["moss_frames"], row["moss_n_vq"])
print(rec["words"][:3], row["caption_general"])

uid contains dots — do not split on the first one. And `audio_key` is not unique across runs: join on uid, or on the triple (voice, gid, audio_key).


Column dictionary

205 columns in `index/`, 83 in `scores/`, 7 in `manifests/shards.parquet`. The complete generated reference — every column, its type and a one-line meaning, including all 40 emo_* and all 114 vn_* with the dimension each code stands for — is in [`COLUMNS.md`](COLUMNS.md). The groups below are the tour.

Identity

columnmeaning
uidunique utterance id. Contains dots — do not split on the first.
originoriginal or repair
variant, voice, gid, audio_keysource-run label, voice id (500), generation-group id, source member name
dataset, shard, src_shardprovenance

audio_key is not unique across runs — join on (voice, gid, audio_key), or use uid.

Audio and text

dur_s (seconds), lang (en/de), text (the generation prompt), text_source (source on every row), n_words, align_status, audio_bytes, moss_frames, moss_n_vq.

All parquet files across both partitions share one identical 205-column schema, so pyarrow.dataset discovery works over the whole index without schema promotion.

moss_frames == round(dur_s × 12.5) holds on 100.000% of rows (the tokenizer rounds; a floor rule matches only ~94.7% and is the wrong invariant).

Vocal bursts

n_bursts, burst_labels (list), burst_starts/burst_ends (lists, seconds), burst_placement (inline / general / none), text_with_bursts — the transcript with (burst:LABEL) inserted between the correct two words. null when placement is not inline.

Scores

emo_* ×40 (Empathic-Insight intensities, 0–7), qual_* ×4, genuineness_0_6, blend_0_10, top_emotion, top_emotion_value, vn_<CODE>_reg ×57 (unclamped regression estimate), vn_<CODE>_bucket ×57 (classification argmax; 0…6, BKGN 0…4, EXPL 0…2).

Model chain: laion/voiceclap-commercial (frozen, 768-d, audio padded/truncated to exactly 30 s at 16 kHz) → laion/voicenet-dimension-predictors-commercial.

VoiceNet reliability, stated plainly. These heads were trained on Gemini-3.5-flash perceptual estimates — not ground truth, not human ratings. Mean held-out MAE 0.744 scale-points, mean Pearson r 0.79, mean bucket accuracy 0.617 (chance ≈0.14 at k=7). Best: EXPL MAE 0.444, R_HEAD 0.478, BKGN 0.493. Worst: S_ASMR 1.316, ARSH 1.257, S_RANT 1.238. `R_MIXD` has r = 0.39 — its labels collapsed and it is close to uninformative. Do not treat the weak axes as measurements. Levels are defined by prose, not a shared numeric anchor, so "high" on one dimension is not commensurate with "high" on another.

scores/ — generation-run scores (origin=original only)

An 83-column table produced during generation on correctly-decoded audio. Join on audio_key. Includes spk_sim, wer, genuineness, blend, reward, rank, is_holdout, in_train_pool, and matrix descriptors emotion, condition, dim, level, block, character, burst_class. There is no equivalent table for `origin=repair`; that run's generation metrics live only in src_meta inside the tars.


Captions

16 wordings per utterance as caption_<template>. caption_general is the pipeline's own output and is identical to caption_clausal — verified on 361,277 rows after the 2026-08-23 regeneration, which is what makes the two code paths (capfix.py for the column, caption_render.py for the templates) safe to keep separate. All 16 wordings share one emotion gate; see below.

templatemean charsregister
minimal34who + top emotion + duration
stage73bracketed stage direction
headline88headline, then detail
narrative179descriptive, third-person
emotive206emotion-first
burst_inline250bursts woven into the delivery clause
casting257casting-call register
technical283recording-and-quality first
directive324second-person instruction to a performer
terse341comma-separated tags
bullets411markdown bullets
tags425key=value, for filtering and conditioning
prose444flowing sentences
clausal453semicolon clauses (= caption_general)
dossier461labelled record, one field per line
verbose629maximal, every clause in full sentences

They vary in clause order, verbosity, register, and burst handling, so a caption-conditioned model sees real paraphrase rather than reshuffling.

What the ordinal words are relative to — read before conditioning

The VoiceNet words are absolute; the emotion clause is corpus-relative. These are different and the difference is deliberate.

VoiceNet ordinal words (vn_<CODE>_bucket → "bright", "measured", "masculine") are absolute. The bucket is the argmax of an ordinal classification head whose levels were each defined by a paragraph of prose in absolute terms, and the renderer maps bucket -> word through a static table using no dataset statistics at all. "bright" means bright on the model's own scale, not "bright for this corpus". style_thr=3.0 and the explicit-content flag are fixed constants too. The caveat is upstream: those anchors are Gemini-3.5-flash estimates, so "absolute" means consistent with the model's training anchors, not calibrated to human ratings.

The emotion clause (reads as …) is relative to the wider corpus, as of 2026-08-23. It previously used an absolute emo_thr=1.0 across 40 Empathic-Insight heads that do not share a scale, which named Interest on 90.6 % of all captions and Bitterness on 0.1 % — it was reporting the scale of the head, not the emotion of the clip. An emotion is now named when it falls in the top 10 % for that emotion (U ≥ 0.90, at most 3, ranked by percentile), against code/capnorm.npz: a pooled tie-aware mid-rank ECDF over 132,833,726 utterances spanning 8 datasets and every language. A clip that clears nothing says `no dominant emotion`. So "reads as sadness" means "unusually sad for this corpus", and Interest now appears on 5.3 % of rows. Full rationale and the gate constants: code/CAPTION_GATE.md.

Consequences worth knowing before you condition on these captions:

  • —The emotion clause is not comparable across corpora the way the VoiceNet words are. It is a percentile within the LAION-TTS corpus, so it is deliberately not re-normalised per dataset — doing that would force every dataset to the same emotional profile by construction.
  • —`no dominant emotion` is a real signal, not a missing value. It fires when a clip is unremarkable on all 40 heads. Dropping those rows biases a training set toward expressive speech.
  • —stats/norm_stats_vprof.json still ships per-dimension and per-emotion distributions computed from this release's own output. Use it if you want normalisation scoped to the voice profiles rather than to the whole corpus; do not mix the two scales in one comparison.

Regenerating or adding a wording

python
import pyarrow.dataset as ds, importlib.util
spec = importlib.util.spec_from_file_location("cr", "code/caption_render.py")
cr = importlib.util.module_from_spec(spec); spec.loader.exec_module(cr)
row = ds.dataset("index", partitioning="hive", format="parquet").head(1).to_pylist()[0]
cr.render(row, template="casting")                        # any of the 16
cr.render(row, template="clausal", polarity="v1_buggy")   # the pre-fix wording

caption_render.py renders from the flat numeric columns and nothing else — no audio, no GPU, no model, ~25 µs per caption. Add your own by putting a function in cr.TEMPLATES; it receives cr.features(row), which has already resolved every ladder, threshold and ranking.


Known issues and corrections

1. A half-speed decode bug was found and fully corrected before release

An earlier annotation pass decoded these files with a routine that flattened stereo frames. The audio is duplicated mono in a 2-channel MP3, so every sample appeared twice and the decoded signal was exactly 2× too long — half speed. Every GPU stage consumed it: the VoiceCLAP embedding (and therefore all 57 VoiceNet dimensions, genuineness, blend), the BUD-E heads (40 emo_*, 4 qual_*), burst detection, forced alignment, and the MOSS tokenizer.

The entire annotation stack was recomputed from the fixed decoder. Verification on the shipped data:

checkresult
dur_s / true duration measured with soundfile1.000000 (mean, median, min, max; n=731 across 24 shards)
moss_frames == round(dur_s × 12.5)100.0000% (621,480 rows across 120 shards)
dur_s / scores.dur (independent generation-run durations, joined on audio_key)1.000000 (mean, median, p1, p99; 40,274 rows)
word-end timestamps beyond dur_s0 / 666 (previously 100% overran)

The MP3 audio was never affected and was not regenerated — the pipeline copies source bytes into the output tar unchanged. Only derived values were wrong, and they have been recomputed.

Why it hid: the internal check moss_frames == floor(dur_s × 12.5) passed on affected rows because both terms were doubled.

2. Caption polarity — corrected

The GEND and BKGN caption ladders ran backwards relative to the data.

  • —`GEND`: vn_GEND_reg correlates +0.842 with vn_R_CHST_reg (chest resonance, a masculine marker), −0.484 with vn_R_HEAD_reg, −0.352 with vn_BRGT_reg. High GEND is masculine — and the model's own documentation agrees (bucket 0 hyper-feminine, top bucket hyper-masculine). The old ladder captioned the most masculine voices "strongly feminine".
  • —`BKGN`: correlates +0.699 with vn_RCQL_reg and +0.29 with qual_background_quality, an independent head. High BKGN is cleaner — the model's documentation agrees (bucket 0 background overwhelms the speech, top acoustically dead). The old ladder captioned clean studio-synthetic audio "very noisy background".

Only the prose inverted — every numeric column was always correct. The pre-fix wording remains reproducible via caption_general_v1_wording and polarity="v1_buggy".

A fourth, independent confirmation of the direction arrived after this text was first written: classifying all 500 voice profiles from the numeric vn_GEND against each profile's own design-spec card_gender agrees on 89.2 % of 379 decided voices. An inverted ladder would score about 11 %.

Superseded claim, kept visible rather than deleted. This section used to say the fix changes only GEND and BKGN terms with zero other differences. That was true of the polarity fix alone and is no longer true of the current captions: the 2026-08-23 regeneration also replaced the emotion clause (issue 6). Across the corpus the two changes together touched 83,146,345 rows for GEND, 79,455,045 for BKGN, and the emotion clause on nearly all of them. Delivery, timbre, speech, affect, style, recording quality, the explicit flag and burst handling remain untouched, and that is asserted per shard.

6. The emotion clause was regenerated on 2026-08-23 — re-read your captions

Every caption in this release was re-rendered. Two things changed and nothing else:

  1. 1.The emotion clause now uses a corpus-percentile gate instead of an absolute threshold (see Captions above). Interest fell from being named on 90.6 % of rows to 5.3 %; all 40 emotions now occur; 17.9 % of rows say no dominant emotion.
  2. 2.`GEND`/`BKGN` polarity (issue 2) is applied to the rendered strings.

The previous string is preserved. caption_general_v1 holds the pre-regeneration caption_general verbatim on every row, so the change is auditable row-by-row and reversible. Note the distinction from the older caption_general_v1_wording, which is a re-render of the pre-fix GEND/BKGN wording over the corrected numbers — not the literal previous string.

If you cached captions before 2026-08-23, rebuild rather than patch. Anything derived from the old text — SFT/DPO prompt strings, retrieval indices, caption-conditioned training sets — carries the inverted GEND/BKGN wording and an Interest-saturated emotion clause. Do not attempt to fix it with find-and-replace: recompute from vn_GEND_bucket / vn_BKGN_bucket and the emotion columns, which is what code/capfix.py does, and which is idempotent.

Determinism note: the ECDF table is kept in float64. An earlier float32 cast collapsed genuinely different percentiles onto one value (emo_Relief 0.99780922730 and emo_Contentment 0.99780920848 both became 0.99780923128), which made tied emotions order nondeterministically. Genuine ties now break by descending percentile, then ascending emotion name — reproducible from the caption alone.

3. There is no Parakeet ASR for these runs

If you want raw ASR with word timestamps: it does not exist here. What exists is MMS_FA forced alignment of the generation prompt text onto the audio — text_source is source on every row. The wer column in scores/ is Whisper large-v3-turbo, computed at generation time. nvidia/parakeet-tdt-0.6b-v3 is used elsewhere in the wider corpus but was never run over the voice profiles. About 1.4% of rows have a non-ok align_status (CTC length failures).

4. Synthetic data with a deliberate stratified bias

Every utterance is model output, not recorded human speech. The 842-condition matrix is a stratified design, not a natural distribution: emotions, voice qualities and edge cases are massively over-represented relative to natural speech. Do not compute population statistics from this corpus. The three vprof_* runs are near-duplicates of the same 500 voices saying the same lines — concatenating them without weighting over-represents this material.

5. The MOSS codes are a re-analysis, not the model's original codes

The audio was generated with MOSS — the model emitted codes and decoded them — but only the decoded MP3 was written. Re-encoding gives codes for the same audio, not the codes the model actually emitted.


Related

Citation

bibtex
@misc{schuhmann_laion_voice_profiles_2026,
  title  = {LAION Voice Profiles — Annotated},
  author = {Schuhmann, Christoph and {LAION}},
  year   = {2026},
  url    = {https://huggingface.co/datasets/laion/laion-voice-profiles-annotated}
}
laion/laion-voice-profiles-annotated · Team Ai