Team Ai
Modelpublic

nityanandmathur/aspect-d-masked-diffusion-tts

sourceHugging Facecc-by-nc-4.0updated 8d agoView on Hugging Face
1likes
Model Card

ASPECT-D: masked-diffusion TTS models for a width × depth × refinement-steps study

This repository holds the trained models behind the paper Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS (Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh; NeurIPS 2026 workshop Diffusion Language Models: Foundations, Efficiency, and Reasoning, poster).

  • —Paper (camera-ready PDF): <https://nityanandmathur.com/aspect-d/assets/paper.pdf>
  • —Project page: <https://nityanandmathur.com/aspect-d/>
  • —Code (training, sampler, evaluation, analysis, inference): <https://github.com/nityanandmathur/aspect-d>
  • —License: weights, generated audio and data-derived files are CC BY-NC 4.0 (non-commercial); the code is MIT. See License.
[!WARNING] This is a research release, not a production TTS system. - It contains 45 small zero-shot text-to-speech models that form one controlled grid: three parameter budgets × five width–depth shapes × three seeds, all trained identically. - They exist so that scaling behaviour can be measured, not to sound as good as possible. All of them use one short, fixed training schedule. - The shallowest, widest shapes (A1, A2, B1) are barely intelligible.

<!-- BEGIN:folders --> 83 checkpoint folders: 45 grid runs, 20 learning-rate sweep runs, and 18 runs under v1.1/ (13 of them evaluated). <!-- END:folders -->

The finding, in plain words

A masked-diffusion TTS model can spend sequential compute in two ways. It owns depth: the number of transformer blocks, fixed when training ends. It rents refinement steps: the number of denoising passes T, chosen per utterance at serving time. We trained one grid over width and depth and then varied T at inference only, which gives a full (width, depth, T) response surface for each metric at no extra training cost.

Refinement does not help every capability equally. It mostly buys intelligibility (word error rate), and buys much less speaker identity (similarity to the prompt voice). The comparison is made against floors that we measure on the same items, not against zero: the error an ASR system makes on the real recordings, and the speaker similarity of a codec round trip of the real recording.

<!-- BEGIN:headline --> Measured against those floors, raising the number of refinement steps from T=1 to T=16 closes 86.2% (95% CI 82.8–89.1%) of the reachable intelligibility range, but only 46.4% (45.0–47.8%) of the reachable identity range: a 1.86× asymmetry (CI 1.82–1.90×), which stays between 1.68× and 1.86× under the monotone reparameterisations of the error that the paper tested. (results/floors.json) <!-- END:headline -->

Retraining the same configurations for 3× and 6× the schedule shrinks this factor (1.84×, 1.36×, 1.23×), but it stays above one: refinement's intelligibility share saturates while its identity share keeps growing. (This repository holds seed 0 of the three 90k-step runs; the other 90k seeds and the 180k-step runs are not uploaded yet.) The paper therefore reports the ordering, not the size of the factor, as the result. What does move identity is training compute and test-time search. At the same number of generator passes (NFE 128), best-of-2 sampling at T=8, with a speaker-verification selector picking the candidate, beats T=16 in speaker similarity. The selector adds a few percent of compute and the gain costs some WER. Longer prompts and per-item speaking-rate matching added little. We could not identify a fixed exchange rate between depth and steps. The pre-registered form in which steps substitute for depth fits worse than one that treats them separately, but neither form fits the data well, so that test stayed undecided. Finally, a large part of the remaining identity gap is lost inside the codec before the model generates anything. Details and every pre-registered verdict, including the failed ones, are in the paper.

Which checkpoint should I use?

  • —To listen to or build on the best model here, use `v1.1/C3_0_90k`. It is the C3 shape (width 768, depth 18) trained for 90,000 steps instead of 30,000, and it has the highest speaker similarity of everything released.
  • —To study scaling, use the 45-run grid (A1_0 … C5_2). Those are the checkpoints whose measurements make up runs.csv, the fits and the paper's main surface.
  • —Use T = 16 steps per codebook level (NFE = 128) unless you are studying T. Quality drops sharply at small T; that drop is the subject of the paper.

<!-- BEGIN:recommend --> | rank by SIM-o | Hub folder | width × depth | steps | WER ↓ | SIM-o ↑ | UTMOS ↑ | |---:|---|---|---:|---:|---:|---:| | 1 | v1.1/C3_0_90k | 768 × 18 | 90,000 | 0.056 | 0.481 | 3.41 | | 2 | v1.1/C1_0_90k | 1152 × 8 | 90,000 | 0.066 | 0.474 | 3.44 | | 3 | v1.1/C5_0_90k | 512 × 36 | 90,000 | 0.053 | 0.464 | 3.37 | | 4 | v1.1/D5_0 | 768 × 38 | 30,000 | 0.105 | 0.416 | 3.04 | | 5 | v1.1/D3_0 | 1088 × 20 | 30,000 | 0.113 | 0.414 | 3.03 |

Top 5 of the 58 released checkpoints that have evaluation scores, ranked by SIM-o at T=16 (single checkpoints, 400 items). The lowest WER of any released checkpoint is v1.1/C5_0_90k (0.053, SIM-o 0.464). Within the 45-run 30k-step grid (means over seeds), C3 has the highest SIM-o (0.392, WER 0.143) and C5 the lowest WER (0.099, SIM-o 0.388).

For reference, the floors measured on the same 400 items: WER 0.0345 (Whisper-large-v3 on the real recordings), SIM-o 0.5554 (a Mimi round trip of the real target, no model in the loop), UTMOS 3.33 (real recordings). <!-- END:recommend -->

WER is the mean per-item word error rate of Whisper-large-v3 (as a fraction, lower is better), SIM-o is the WavLM-large speaker-verification cosine similarity to the original prompt recording (higher is better), and UTMOS is UTMOS22-strong predicted naturalness (1–5). See Evaluation.

Quick start

The models are plain PyTorch modules defined in the code repository. Inference needs that repository, espeak-ng (for phonemes), and the Mimi codec (`kyutai/mimi`, downloaded automatically through transformers). The commands below were run on macOS (CPU and Apple MPS) with Python 3.11. synthesize.py has not yet been run on CUDA; the paper's own evaluation ran the same sampler on NVIDIA B200 GPUs.

bash
git clone https://github.com/nityanandmathur/aspect-d && cd aspect-d
brew install espeak-ng                 # Linux: sudo apt-get install espeak-ng
uv venv --python 3.11 .venv
uv pip install --python .venv/bin/python torch transformers safetensors huggingface_hub \
    phonemizer soundfile librosa numpy pandas

.venv/bin/python src/synthesize.py --run v1.1/C3_0_90k \
    --text "Our goal is gonna be to win a whole round, buddy. Get that yellow, get that yellow!" \
    --prompt-wav examples/prompts/it0001_prompt.flac \
    --prompt-text "Oh, here we go! Get me close." \
    --steps 16 --seed 0 --device auto --out out.wav

--run takes any checkpoint folder of this repository (C3_0, v1.1/D3_0, …). The script prints a JSON record with the number of frames, NFE, the count of phonemes unknown to the model, and timings. The prompt transcript is required and must match the prompt audio. examples/prompts/ in the code repository holds the two real Emilia prompt recordings used here, with their transcripts.

From Python (run from the root of the code checkout):

python
import sys; sys.path.insert(0, "src")
import soundfile as sf
from synthesize import SR, load_tts, synthesize

tts = load_tts(run="v1.1/C3_0_90k", device="auto")      # or load_tts(checkpoint="path/to/folder")
wav, info = synthesize(tts, text="Hello there, this is a test.",
                       prompt_wav="examples/prompts/it0000_prompt.flac",
                       prompt_text="Yep. So, shouldn't you just get network plus so that you get that?",
                       steps=16, seed=0)
sf.write("out.wav", wav, SR)                             # 24 kHz mono, generated target only
print(info["target_seconds"], info["nfe"], info["oov_phonemes"])

Loading only the network (no phonemizer, no codec), for example to inspect weights:

python
import json, sys
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
sys.path.insert(0, "src")                                # aspect-d checkout
from model import AspectD

repo, run = "nityanandmathur/aspect-d-masked-diffusion-tts", "C3_0"
cfg = json.load(open(hf_hub_download(repo, f"{run}/config.json")))
vocab = json.load(open(hf_hub_download(repo, "phone_vocab.json")))["vocab"]
model = AspectD(cfg["width"], cfg["depth"], cfg["heads"], n_phonemes=len(vocab))
model.load_state_dict(load_file(hf_hub_download(repo, f"{run}/model.safetensors")), strict=True)
model.eval()        # weights are stored in bfloat16 and copied into float32 parameters
assert model.nonembed_params() == cfg["nonembed_params"]
assert model.total_params() == cfg["total_params"]

What synthesize.py does, step by step, all through the same functions that produced the paper's numbers:

  1. 1.Phonemize the prompt transcript and the target text with espeak-ng (en-us, stress marks, punctuation kept) and map them to phone_vocab.json. The visible phoneme sequence is the prompt transcript, a word break, then the target text.
  2. 2.Convert the prompt to mono 24 kHz, peak-normalise it, and encode it with Mimi (8 codebooks). Prompts longer than 3.5 s are cut to their first 3.5 s.
  3. 3.Fix the output length before generation: number of characters × the corpus rate in dataset.json × 12.5 frames per second, capped at 15 s.
  4. 4.Sample with the frozen MaskGIT-style sampler: codebook level by level, T steps per level (NFE = 8T), cosine unmasking schedule, temperature 1.0, Gumbel confidence noise annealed from 1 to 0, no classifier-free guidance.
  5. 5.Decode prompt + target with Mimi and return only the target.

Single calls do not reproduce the paper's evaluation audio bit for bit: the paper sampled 50 items per batch (the noise depends on the batch shape) in bfloat16 on NVIDIA B200 GPUs. The code path is the same. The same seed gives the same audio on the same device; CPU and MPS give different audio.

What is in this repository

Checkpoint folders. Every folder holds:

  • —model.safetensors: the weights in bfloat16, without optimizer state. Keys match src/model.py exactly.
  • —config.json: shape (width, depth, heads), parameter counts, base learning rate, training steps, final validation loss and training GPU-hours.
  • —run.json: the full training record, including the validation-loss history every 1,000 steps.
folderswhat they are
A1_0 … C5_2 (45)The main grid: configurations A1–A5, B1–B5, C1–C5 × seeds 0, 1, 2, all 30,000 steps. Budgets A, B and C target 20M, 50M and 125M non-embedding parameters; within a budget, shape 1 is the widest and shape 5 the deepest. These are the models measured in runs.csv.
sweep_g1_w256_lr*, sweep_g1_w640_lr* (10)μP width-transfer check (protocol gate G1), run before the grid: 3,000-step proxies at depth 12 and width 256 or 640, each at base LR 0.001, 0.002, 0.004, 0.008, 0.016. The gate passes if the best LR agrees within a factor of 2.
sweep_g1b_d4_lr*, sweep_g1b_d24_lr* (10)μP depth-transfer check (gate G1b): the same five LRs at width 384 and depth 4 or 24.
v1.1/D1_0 … v1.1/D5_1 (10)Extension: a fourth, larger budget D (nominal 285M non-embedding parameters in grid.json; the paper's 276M is the mean actual count of these ten checkpoints), five shapes × two seeds, 30,000 steps, base LR 0.002. Used for the parameter axis and the scale-persistence test.
v1.1/C1_0_90k, v1.1/C3_0_90k, v1.1/C5_0_90k (3)Extension: seed 0 of three C-budget shapes trained from scratch for 90,000 steps (3× the grid schedule), as a training-compute control.
v1.1/lrsweep_D3_0p001, _0p002, _0p008, _0p016 (4)Extension: 3,000-step base-LR sweep for the D budget at shape D3. It was the pre-registered remedy after the D-budget sanity gate (G1-D) failed; its best LR, 0.002, was used for all D runs.
v1.1/proxy_D3_g1d (1)Extension: the 3,000-step G1-D sanity proxy at D3 with base LR 0.004, which is also the 0.004 point of that sweep.

The sweep and proxy checkpoints were only ever compared by validation loss; they were never used to synthesize speech.

Other files.

filecontents
runs.csvThe measured surface of the 45-run grid: one row per (config, seed, T) for T ∈ {1, 2, 4, 8, 16}, 225 rows. Columns: shape (config, seed, budget, width, depth, heads, n_nonembed, n_total, lr, ckpt_step); T and nfe (= 8T); val_loss; wer (mean per-item WER over all 400 items, degenerate outputs included), wer_se, wer_corpus (corpus-level WER); sim (SIM-o), sim_se, sim_model; utmos, utmos_se; degen_rate (share of outputs flagged by the frozen degenerate-output rule) and crash_rate; wer_nondegen, sim_nondegen (means over non-degenerate items); err_wer = WER, err_sim = 1 − SIM-o, err_ut = (5 − UTMOS)/4 (the error transforms the fits use); c_layer_ms (measured time per layer at that width) and latency_ms = 8 · T · depth · c_layer_ms; train_gpu_hours, train_wall_s, synth_gpu_hours; n_items.
fits.jsonThe pre-registered fits of that surface: the shape test at fixed T (part_a), the separable and substitution scaling forms with their exponents (part_b), gate G4, the small-to-large prediction test (hd4), saturation, degenerate rate by T, the 2,000-replicate run-level bootstrap, and the decision values (Δτ, κ, τ per metric). Caution: Δτ changes sign under monotone reparameterisations of the error and κ sits at its bound (paper §4.3–4.4); the paper reports the gap-closed ratio, not Δτ, as its result.
grid.jsonThe canonical design: backbone and parameter formula, codec, conditioning, training recipe, sampler, budgets, all 20 shapes (including D), seeds, evaluation models.
dataset.jsonTraining and evaluation data statistics, including sec_per_char, the corpus character rate that sets the output length.
phone_vocab.jsonThe phoneme vocabulary (159 symbols, ids of the phoneme embedding). Required for inference.
protocol.htmlThe pre-registered protocol: gates, thresholds, statistics and deliverables. The current copy, with later annotations marked as such, is the code repository's docs/protocol.html.
results/extension_runs.csvScores of the 13 evaluated v1.1/ checkpoints at every T that was synthesized (T = 1 to 64), copied from the code repository's committed results/runs-v1.1/*/synth_T*/scores.json; each row names its source file.
results/floors.jsonThe measured floors and the headline numbers above, with the code-repository file behind each value.
results/compute.jsonTraining GPU-hours of the extension runs whose records are committed in the code repository (results/runs-v1.1/, whose weights are here as v1.1/, and results/runs-v1.3/, whose weights are not), with the device.
scripts/make_card_tables.pyWrites every number and table in this card from the files above.

Checkpoints of the main grid

<!-- BEGIN:gridtable --> | config | budget | width | depth | heads | non-emb. params | total params | Hub folders | final val loss | WER | SIM-o | UTMOS | |---|---|---:|---:|---:|---:|---:|---|---:|---:|---:|---:| | A1 | A | 640 | 4 | 10 | 19.7M | 40.8M | `A10, A11`, `A12 | 4.638 | 0.539 | 0.327 | 2.84 | | A2 | A | 448 | 8 | 7 | 19.3M | 34.0M | A20`, `A21, A22` | 4.582 | 0.284 | 0.348 | 2.91 | | A3 | A | 384 | 12 | 6 | 21.2M | 33.9M | `A30, A31`, `A32 | 4.566 | 0.159 | 0.360 | 2.96 | | A4 | A | 320 | 18 | 5 | 22.1M | 32.7M | A40`, `A41, A42` | 4.574 | 0.136 | 0.359 | 2.94 | | A5 | A | 256 | 24 | 4 | 18.9M | 27.3M | `A50, A51`, `A52 | 4.611 | 0.153 | 0.351 | 2.92 | | B1 | B | 832 | 6 | 13 | 49.9M | 77.3M | B10`, `B11, B12` | 4.551 | 0.278 | 0.361 | 2.93 | | B2 | B | 640 | 10 | 10 | 49.2M | 70.2M | `B20, B21`, `B22 | 4.505 | 0.160 | 0.381 | 2.98 | | B3 | B | 512 | 16 | 8 | 50.3M | 67.2M | B30`, `B31, B32` | 4.497 | 0.119 | 0.388 | 3.00 | | B4 | B | 448 | 22 | 7 | 53.0M | 67.8M | `B40, B41`, `B42 | 4.511 | 0.120 | 0.382 | 2.99 | | B5 | B | 384 | 30 | 6 | 53.1M | 65.8M | B50`, `B51, B52` | 4.520 | 0.115 | 0.375 | 2.98 | | C1 | C | 1152 | 8 | 18 | 127.4M | 165.4M | `C10, C11`, `C12 | 4.504 | 0.171 | 0.373 | 2.98 | | C2 | C | 960 | 12 | 15 | 132.7M | 164.4M | C20`, `C21, C22` | 4.504 | 0.189 | 0.373 | 2.97 | | C3 | C | 768 | 18 | 12 | 127.4M | 152.7M | `C30, C31`, `C32 | 4.488 | 0.143 | 0.392 | 2.98 | | C4 | C | 640 | 26 | 10 | 127.8M | 148.9M | C40`, `C41, C42` | 4.489 | 0.134 | 0.390 | 3.01 | | C5 | C | 512 | 36 | 8 | 113.3M | 130.2M | `C50, C51`, `C52` | 4.487 | 0.099 | 0.388 | 3.02 |

15 configurations, 45 checkpoints. WER, SIM-o and UTMOS are at T=16 (NFE=128) on the 400 evaluation items, averaged over the seeds listed; final val loss is the mean over the same seeds (masked cross-entropy at a fixed mask grid, comparable across runs). All checkpoints are at step 30,000. Source: runs.csv (T=1,2,4,8,16 per run) and each folder's config.json. <!-- END:grid_table -->

Non-embedding parameters are those of the transformer blocks and the final norm. They equal 12·depth·width² (the attention and MLP matrices) plus the RMSNorm gains, so they differ slightly from the nominal counts in grid.json. Total parameters add the phoneme and codec embeddings and the eight output heads.

Extension checkpoints (v1.1/)

<!-- BEGIN:exttable --> | Hub folder | what it is | width | depth | non-emb. params | steps | base LR | final val loss | WER | SIM-o | UTMOS | |---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:| | `v1.1/C1090k` | C-budget shape trained from scratch for 90k steps | 1152 | 8 | 127.4M | 90,000 | 0.004 | 4.091 | 0.066 | 0.474 | 3.44 | | `v1.1/C3090k` | C-budget shape trained from scratch for 90k steps | 768 | 18 | 127.4M | 90,000 | 0.004 | 4.119 | 0.056 | 0.481 | 3.41 | | `v1.1/C5090k` | C-budget shape trained from scratch for 90k steps | 512 | 36 | 113.3M | 90,000 | 0.004 | 4.179 | 0.053 | 0.464 | 3.37 | | `v1.1/D10 | D-budget shape, 30k steps | 1536 | 10 | 283.1M | 30,000 | 0.002 | 4.415 | 0.137 | 0.407 | 3.03 | | v1.1/D11` | D-budget shape, 30k steps | 1536 | 10 | 283.1M | 30,000 | 0.002 | 4.413 | 0.148 | 0.394 | 3.05 | | `v1.1/D20 | D-budget shape, 30k steps | 1280 | 14 | 275.3M | 30,000 | 0.002 | 4.401 | 0.100 | 0.403 | 3.03 | | v1.1/D21` | D-budget shape, 30k steps | 1280 | 14 | 275.3M | 30,000 | 0.002 | 4.396 | 0.095 | 0.409 | 3.07 | | `v1.1/D30 | D-budget shape, 30k steps | 1088 | 20 | 284.1M | 30,000 | 0.002 | 4.414 | 0.113 | 0.414 | 3.03 | | v1.1/D31` | D-budget shape, 30k steps | 1088 | 20 | 284.1M | 30,000 | 0.002 | 4.413 | 0.119 | 0.409 | 3.05 | | `v1.1/D40 | D-budget shape, 30k steps | 896 | 28 | 269.8M | 30,000 | 0.002 | 4.414 | 0.121 | 0.413 | 3.03 | | v1.1/D41` | D-budget shape, 30k steps | 896 | 28 | 269.8M | 30,000 | 0.002 | 4.418 | 0.151 | 0.409 | 3.03 | | `v1.1/D50 | D-budget shape, 30k steps | 768 | 38 | 269.0M | 30,000 | 0.002 | 4.401 | 0.105 | 0.416 | 3.04 | | v1.1/D5_1` | D-budget shape, 30k steps | 768 | 38 | 269.0M | 30,000 | 0.002 | 4.406 | 0.093 | 0.410 | 3.04 |

Single checkpoints at T=16, 400 items. Source: results/extension_runs.csv (every T that was scored, T=1 to 64; rows with n_items=200 are the 200-item extended-T subset) and each folder's config.json. <!-- END:ext_table -->

Model details

  • —Backbone. A bidirectional (non-causal) pre-LN transformer with RMSNorm, rotary position embeddings, a GELU MLP of ratio 4, head dimension 64 and no weight tying. The sequence is [phonemes | audio frames]; attention is full and bidirectional.
  • —Tokens. Mimi codec tokens at 12.5 frames per second, using the first 8 of Mimi's residual codebooks (2,048 entries each, plus MASK and PAD). One embedding table per codebook is summed at the input, and one linear head per codebook reads out.
  • —Conditioning. espeak-ng phonemes of the prompt transcript and the target text, and the codec tokens of a roughly 3-second prompt at all 8 levels, are always visible. There is no speaker encoder: continuing the prompt is the only way identity reaches the output. The output length is fixed before generation from the text length. There is no classifier-free guidance, in training or at inference.
  • —Training objective. SoundStorm-style coarse-to-fine masking. For each example a codebook level l is drawn uniformly from 1–8 and a mask ratio t uniformly from (0, 1]; levels below l stay visible, cells of level l are masked independently with probability t, and levels above l are entirely masked. The loss is cross-entropy on the masked cells of level l.
  • —Sampler. MaskGIT-style confidence decoding, frozen for every model and every T: level by level, exactly T steps per level (NFE = 8T), a cosine schedule for how many cells stay masked, tokens drawn at temperature 1.0, confidence = log-probability plus Gumbel noise annealed from 1 to 0, committed cells stay committed.
  • —Output. Mimi tokens, decoded by the Mimi decoder to 24 kHz mono audio. Mimi is not part of this repository; it is loaded from `kyutai/mimi`.

Training details

<!-- BEGIN:training --> | | | |---|---| | training data | 2,000 h of English speech from Emilia (Emilia/EN shards; 71 shards scanned), 792,064 clips, 64,680 speakers, speaker-stratified selection | | held-out speakers | 200, never seen in training; the evaluation items come from these | | validation set | 2,000 clips from 2,000 speakers (the source of every validation loss in this card) | | steps, batch | 30,000 steps at an effective batch of 256 sequences (prompt 3 s + target up to 15 s) | | optimiser | AdamW, betas (0.9, 0.95), weight decay 0.1, gradient clip 1.0, 600 warm-up steps, cosine to 10% of peak, bf16 | | μP | base width 256, base LR 0.004 for budgets A–C (from 5-point sweeps over 0.001, 0.002, 0.004, 0.008, 0.016), residual scale 1/sqrt(2*depth) | | length rule | 0.0602 s per character (median over 10,000 training clips) | | compute | 162.5 GPU-hours of training for the 45 grid checkpoints (sum of train_gpu_hours in runs.csv), the total the paper's App. A reports. The other figures come from the run records, not the paper: with the 20 learning-rate sweep runs (6.9, their config.json) and grid synthesis (2.4, runs.csv), 171.8 GPU-hours for the main study. The 18 v1.1/ checkpoints took 89.2 GPU-hours of training (86.2 for the 13 evaluated ones); the committed run records of the extension training runs sum to 169.3 GPU-hours, including 8 runs whose weights are not here (results/compute.json). All values are sums of job wall-clock time on NVIDIA B200 GPUs; small jobs often shared a GPU, so they overstate exclusive GPU time. | <!-- END:training -->

Every shape uses the same steps, batch, schedule, masking random stream, data order and evaluation cadence; only width and depth change within a budget. The grid's design is in grid.json and the code repository's docs/protocol.html.

Evaluation

<!-- BEGIN:eval --> 400 cross-sentence zero-shot items from 174 of the 200 held-out speakers: the prompt is one complete real clip of 2.5–3.5 s with its transcript, and the target is a different utterance (4–15 s) of the same speaker, so copying the prompt's content scores nothing. Items were drawn from 6,810 candidates, keeping only those whose real recording Whisper-large-v3 transcribes with WER ≤ 0.25 (6,520 qualified; this removes mislabelled non-English clips in Emilia's EN split); the selected items' mean ground-truth WER is 0.0345. Predicted target durations span 4.0–14.9 s. <!-- END:eval -->

  • —Intelligibility: WER of Whisper-large-v3 (greedy decoding, Whisper English text normaliser) against the target text.
  • —Identity: SIM-o, the cosine similarity between WavLM-large speaker-verification embeddings (the UniSpeech checkpoint used by seed-tts-eval) of the output and of the original prompt waveform.
  • —Naturalness: UTMOS22-strong (a tertiary, predicted metric).
  • —Reliability: a frozen degenerate-output rule (empty transcript, under 30% of the target character count, or a 4-gram repeated 6 or more times in a row). Primary means include degenerate items.

Every gain in the paper is read against floors measured on the same 400 items through the same metric stack:

<!-- BEGIN:floors --> | reference point | value | what it is | |---|---:|---| | WER floor | 0.0345 | mean per-item WER of Whisper-large-v3 on the REAL target recordings of the 400 eval items | | SIM-o ceiling | 0.5554 | mean SIM-o (WavLM-large SV vs the original prompt) of the real target after a Mimi encode/decode round trip, no model in the loop | | SIM-o of real audio | 0.6784 | mean SIM-o of the real target recording itself (no round trip) | | UTMOS of real audio | 3.33 | mean UTMOS22-strong of the real target recordings |

Source: results/floors.json, which records the code-repository file behind each value. <!-- END:floors -->

Limitations

  • —Small, with a short training schedule. The grid tops out at 133M non-embedding parameters and 30,000 steps, and the largest model here has under 300M non-embedding parameters. Speech quality is well below production systems; the shallowest shapes (A1, A2, B1) are barely intelligible even at T=16; see the grid table.
  • —One language, one codec, one corpus. English only (Emilia-EN). Audio quality is capped by 8-codebook Mimi at 12.5 Hz.
  • —No classifier-free guidance, nothing tuned. Temperature 1.0 and the frozen sampler for every model. The models were built to be compared with each other, not tuned for quality.
  • —Length comes from a corpus median. The output duration is set from the character count and the corpus-median rate in dataset.json. There is no duration predictor, and the rate does not adapt to the speaker. Text longer than about 15 s of speech is squeezed into 15 s; synthesize long text sentence by sentence.
  • —The prompt transcript is required and must match the prompt audio exactly; otherwise the model tends to speak the leftover words at the start of the output. Noisy or reverberant prompts are continued as they are.
  • —espeak-ng version. The training phonemes came from espeak-ng 1.51. Other versions can phonemize some words differently. Symbols never seen in training become <unk> (synthesize.py reports how many), but a different in-vocabulary phone choice cannot be detected.
  • —Metrics are model-based. WER, SIM-o and UTMOS come from Whisper, WavLM and UTMOS; the paper reports how its conclusions depend on the choice of scorer and coordinate.
  • —No watermarking is applied to generated audio.

Intended use and misuse

These models are for research: scaling behaviour, test-time compute and masked-diffusion generation of speech, and reproducing or extending the paper's analyses.

They clone a voice from about 3 seconds of audio. Please:

  • —only clone voices whose owners have given consent;
  • —do not impersonate real people, deceive listeners, commit fraud, or try to bypass voice-based authentication;
  • —say clearly that generated audio is synthetic when you share it;
  • —do not use the models commercially (see the license) or as a production TTS system; they are research artifacts and were not hardened against misuse.

License

The weights, config.json/run.json records, audio generated with the models and the data-derived metadata (phone_vocab.json, dataset.json) are released under [CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/legalcode), the most restrictive license among the inputs: the training data, Emilia, is CC BY-NC 4.0. The Mimi codec is CC BY 4.0. The code in the GitHub repository is MIT. The full statement, with the reasoning, is `MODEL_LICENSE.md` in the code repository.

The measurement and design files (runs.csv, fits.json, grid.json, results/, protocol.html) are also CC BY-NC 4.0. scripts/make_card_tables.py is code and, like the GitHub repository, MIT-licensed.

The two prompt clips that the quick start uses, in the code repository's examples/prompts/, are real Emilia recordings. Emilia does not own the copyright of its audio; it stays with the owners of the original recordings. The clips are therefore shared only under Emilia's non-commercial research terms, and a rights holder who wants one removed can open a discussion on this Hub repository or an issue on the code repository.

If you use these models or audio, please also credit Emilia (He et al., Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation, IEEE SLT 2024, <https://doi.org/10.1109/SLT61566.2024.10832365>) and Mimi (Kyutai; Défossez et al., Moshi: a speech-text foundation model for real-time dialogue, <https://arxiv.org/abs/2410.00037>).

Citation

bibtex
@inproceedings{mathur2026refinement,
  title     = {Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion {TTS}},
  author    = {Mathur, Nityanand and Sayed, Hamees and Singh, Ayush Pratap},
  booktitle = {NeurIPS 2026 Workshop on Diffusion Language Models: Foundations, Efficiency, and Reasoning (DiffuLM)},
  year      = {2026},
  url       = {https://openreview.net/forum?id=E659lrDKOx},
  note      = {Poster}
}