cloud0day3/antalia-eval
Antalia evaluation suites and training-data manifests Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the aggregate statistics used by the best-of-N selector. The speaker's own… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-eval.
Antalia evaluation suites and training-data manifests
Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the aggregate statistics used by the best-of-N selector.
The speaker's own recordings are published separately as antalia-voice-corpus (5.008 h, CC-BY-4.0).
Code: https://github.com/0daycloud/antalia · Paper: https://github.com/0daycloud/antalia/blob/main/paper/main.pdf
Contents
evaluation/ — prompt suites (CC-BY-4.0, written by the authors)
Fields: id, category, text. Prompts were deduplicated against every training manifest.
manifests/ — public-corpus filter manifests
Reproduce exactly which upstream clips passed our gates. Audio is not re-hosted; obtain it from Mozilla Common Voice 26.0 (Turkish) and Google FLEURS (tr_tr) under their own terms.
Per record: clip_id, source_file (upstream file name, e.g. common_voice_tr_29908728.mp3), transcript, normalized_transcript, duration_seconds, sample_rate_hz, sha256 of the upstream audio file, acoustic QA (common_voice_acoustics: peak, RMS dBFS, clipping ratio, active frame ratio, estimated SNR), split, license id, attribution.
Removed on purpose: speaker identifiers, age, gender, vote counts, and local paths. The Mozilla Data Collective terms ask users not to attempt speaker identification; these manifests add no capability beyond the upstream dataset. The foundation model's training subset (59,593 Common Voice + 1,876 FLEURS clips, 67.55 h) was drawn from the accepted train split with per-speaker caps applied on upstream client_id; the cap procedure is in the code repository (src/turkish_tts/common_voice_prepare.py).
listening/ — native-listener CMOS protocol and results
cmos-protocol.md (≥3-listener protocol as designed), cmos-v1-instructions.md (Turkish listener sheet), cmos-v1-trials.csv (30 trials: 15 A/B gated-vs-plain, 12 anchored synth-vs-real, 3 catch), cmos-v1-key.json (which system was A/B), cmos-v1-results-listener-1.csv (the only completed session: one native listener, the project owner), cmos-v1-results.md (analysis: pooled CMOS −0.667 ± 0.653; anchored −1.833 ± 0.757; same-person 0/12). The trial audio is not included because half of the anchored pairs are real recordings of the voice actor.
selection/ — best-of-N selector inputs
inference-recipe.json (champion sampling recipe), prosody-presets.json (six style targets), timbre-profile.json (level-normalized 11-band spectrum mean/std of 13 real recordings), envelope-stats.json (F0/energy/rate statistics). These are aggregate statistics; they do not contain audio.
Licenses
Prompt suites, protocol documents, presets, and statistics: CC-BY-4.0 (attribution: the Antalia authors). Manifests describe third-party data: Common Voice clips are CC0-1.0 under the Mozilla Data Collective terms; FLEURS is CC-BY-4.0 (Google). The manifests themselves are released CC-BY-4.0.
Citation
@misc{antalia1_2026,
title = {Antalia 1: An Open Turkish Text-to-Speech Model from a Rights-Clean Pipeline},
author = {Saygili, Sezgin and Kaplaner, Emre and Ozgul, Oncel and Koktas, Fikri San},
year = {2026},
note = {Technical report},
url = {https://huggingface.co/datasets/cloud0day3/antalia-eval}
}Contact: sezgin@patientdesk.ai
