Team Ai
Apppublic

HumeAI/voice-replication-leaderboard

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
10likes
App README

Voice Replication Leaderboard

How well text-to-speech models reproduce a specific person's voice from a short reference recording, rated by paid blind human raters.

Thirteen models replicate the same 25 reference voices across 7 delivery-neutral prompts, three raters per clip. One board tab, introduced by a short blurb above the table, with four subtabs: Overall first, then the three reference types (Standard, Emotional, Accent). Every subtab carries the same four columns, so the categories compare like with like.

  • —Same speaker (the ranking column): does the output sound like the reference person, 1 to 5
  • —Quality: audio quality, independent of the match
  • —Naturalness: how human the delivery sounds
  • —Objective speaker similarity: how close the output's speaker embedding sits to the reference's, 0 to 1. An automatic check on the human ratings. The footnote under the table names the model (TitaNet). It is not a probability and not a percentage, so the column carries no unit.

One human study per model. Round 2 (study 6d81f631) rated VoxCPM2 and LongCat beside two round-1 anchors, fishaudio/s2-pro and cartesia/sonic-3.5. Round 3 (study 296cf503) rated google/gemini-3.8-flash and google/gemini-3.8-flash-lite beside cartesia/sonic-3.5 and openbmb/VoxCPM2. The anchors link the studies' scales; their re-ratings are a calibration check and do not enter the board; each model's numbers come from exactly one rater pool (HUMAN_STUDY in build_boards.py).

Overall weights the three reference categories equally. Each category contributes one third of a model's Overall score, on every column. A plain mean over the 175 samples would let the 15 accent references carry 60% of it against 5 standard and 5 emotional. Category subtabs are plain within-category means and are unaffected.

Only the Overall subtab collapses figures under Detailed view, and it carries two. The first shows the category spread per model, which a single number hides. The second shows where the objective check disagrees with the raters: one bar per model, its human rank minus its objective rank, so no bar means the two agree (Spearman rho = +0.70 over 9 models). The category subtabs carry no figure on purpose: a per-reference chart would publish the grid's reference voices.

Sibling of the Real World VoiceEQ Benchmark. Shares the app machinery and board-JSON schema of the other Hume leaderboard Spaces, on its own private dataset.

Naming: the humeval eval (voice_clone), this repo folder, the board ids and every filename stay "cloning". That is the eval name in the humeval DB and run identity depends on it, so renaming it would orphan the ratings. Only the Hugging Face repos and what a viewer reads on the Space say "replication".

Data

Loaded at runtime from the dataset configured on the Space:

  • —Variable LEADERBOARD_DATASET = HumeAI/voice-replication-leaderboard-data
  • —Secret HF_TOKEN = a read token (the dataset is private)

build_boards.py in this folder produces the board JSON from the humeval prod DB, on canonical runs only, with no pinned study:

bash
HUMEVAL_TARGET=prod uv run python leaderboard/voice-cloning/build_boards.py

A model joins the board only when its reference set matches the current 25-voice grid. higgs-v2 is canonical but still sits on the pre-2026-08-27 grid, so the builder drops it rather than compare it against models that ran the new references.

Local preview

bash
pip install -r requirements.txt
python app.py   # with LEADERBOARD_DATASET unset, reads ./data/

Deployment

Live repos: Space HumeAI/voice-replication-leaderboard, dataset HumeAI/voice-replication-leaderboard-data. Both are private. Deploy with the hf CLI (pip install -U huggingface_hub, then hf auth login with write access to HumeAI), running the commands from this folder. Upload the dataset first: app.py reads the board JSON at startup and bakes it into the Gradio UI, so a Space upload that restarts before the data lands boots against the old JSON. Re-runs are always safe, since unchanged files are skipped.

First time

bash
hf repo create HumeAI/voice-replication-leaderboard-data --repo-type dataset --private
hf repo create HumeAI/voice-replication-leaderboard --repo-type space --space_sdk gradio --private
hf upload HumeAI/voice-replication-leaderboard-data data . --repo-type dataset
hf upload HumeAI/voice-replication-leaderboard . . --repo-type space --exclude "data/*"

Then, in the Space's Settings → Variables and secrets, add the LEADERBOARD_DATASET variable and the HF_TOKEN secret listed above.

Updates

bash
hf upload HumeAI/voice-replication-leaderboard-data data . --repo-type dataset
hf upload HumeAI/voice-replication-leaderboard . . --repo-type space --exclude "data/*"

Pushing to the Space rebuilds and restarts it. A dataset-only change needs an explicit restart, because the running app never re-reads the dataset:

bash
curl -X POST -H "Authorization: Bearer $HF_TOKEN" \
  https://huggingface.co/api/spaces/HumeAI/voice-replication-leaderboard/restart

The audio behind samples.json (data/samples/**) is gitignored and lives only in the dataset. hf upload only adds and overwrites, so a fresh clone with no local clips still deploys safely.

Examples are one flat list. samples.json carries Standard and Accent examples with no factor field, so the app renders them under a single "Examples" heading without category tabs. The Emotional examples were removed; the Emotional subtab keeps its scores.