Team Ai
Datasetpublic

cloud0day3/antalia-eval

Antalia evaluation suites and training-data manifests Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the aggregate statistics used by the best-of-N selector. The speaker's own… See the full description on the dataset page: https://huggingface.co/datasets/cloud0day3/antalia-eval.

sourceHugging Facecc-by-4.0updated 26d agoView on Hugging Face
1likes212downloads
Dataset Card

Antalia evaluation suites and training-data manifests

Companion data for Antalia 1 and Antalia 1 Foundation, an open Turkish text-to-speech release whose development is discontinued. No audio is included. The repository contains the prompt suites we report numbers on, the manifests that reproduce our public-corpus filtering, the native-listening (CMOS) protocol with its raw single-listener results, and the aggregate statistics used by the best-of-N selector.

The speaker's own recordings are published separately as antalia-voice-corpus (5.008 h, CC-BY-4.0).

Code: https://github.com/0daycloud/antalia · Paper: https://github.com/0daycloud/antalia/blob/main/paper/main.pdf

Contents

evaluation/ — prompt suites (CC-BY-4.0, written by the authors)

FilePromptsNotes
turkish-v2.jsonl12010 categories × 12: general, voiceagent, namesplaces, numeric, foreignabbreviations, adversarialnormalization, questionsconfirmations, acknowledgement, emotionalstyle, long_form. All reported CER/WER/similarity numbers use this suite.
turkish-v1.jsonl39Earlier suite used for the FastPitch baseline failure analysis. ack-002 collides with two Common Voice sentences; neither was in any training manifest.
turkish-pronunciation-v1.jsonl39Pronunciation stress set (foreign names, abbreviations, numerals).

Fields: id, category, text. Prompts were deduplicated against every training manifest.

manifests/ — public-corpus filter manifests

Reproduce exactly which upstream clips passed our gates. Audio is not re-hosted; obtain it from Mozilla Common Voice 26.0 (Turkish) and Google FLEURS (tr_tr) under their own terms.

FileRecordsMeaning
common-voice-26.0-tr-filtered.train.jsonl101,622Accepted clips, speaker-disjoint train split (110.2 h)
common-voice-26.0-tr-filtered.validation.jsonl5,014Accepted, validation split (5.1 h)
common-voice-26.0-tr-filtered.test.jsonl4,824Accepted, test split (5.1 h)
common-voice-26.0-tr-filtered.rejected.jsonl8,947Rejected clips with quality_filter_reasons
common-voice-26.0-tr-filtered.report.json—Aggregate counts and rejection reasons
fleurs-tr.jsonl3,607FLEURS Turkish clips (dataset revision 70bb2e84…)

Per record: clip_id, source_file (upstream file name, e.g. common_voice_tr_29908728.mp3), transcript, normalized_transcript, duration_seconds, sample_rate_hz, sha256 of the upstream audio file, acoustic QA (common_voice_acoustics: peak, RMS dBFS, clipping ratio, active frame ratio, estimated SNR), split, license id, attribution.

Removed on purpose: speaker identifiers, age, gender, vote counts, and local paths. The Mozilla Data Collective terms ask users not to attempt speaker identification; these manifests add no capability beyond the upstream dataset. The foundation model's training subset (59,593 Common Voice + 1,876 FLEURS clips, 67.55 h) was drawn from the accepted train split with per-speaker caps applied on upstream client_id; the cap procedure is in the code repository (src/turkish_tts/common_voice_prepare.py).

listening/ — native-listener CMOS protocol and results

cmos-protocol.md (≥3-listener protocol as designed), cmos-v1-instructions.md (Turkish listener sheet), cmos-v1-trials.csv (30 trials: 15 A/B gated-vs-plain, 12 anchored synth-vs-real, 3 catch), cmos-v1-key.json (which system was A/B), cmos-v1-results-listener-1.csv (the only completed session: one native listener, the project owner), cmos-v1-results.md (analysis: pooled CMOS −0.667 ± 0.653; anchored −1.833 ± 0.757; same-person 0/12). The trial audio is not included because half of the anchored pairs are real recordings of the voice actor.

selection/ — best-of-N selector inputs

inference-recipe.json (champion sampling recipe), prosody-presets.json (six style targets), timbre-profile.json (level-normalized 11-band spectrum mean/std of 13 real recordings), envelope-stats.json (F0/energy/rate statistics). These are aggregate statistics; they do not contain audio.

Licenses

Prompt suites, protocol documents, presets, and statistics: CC-BY-4.0 (attribution: the Antalia authors). Manifests describe third-party data: Common Voice clips are CC0-1.0 under the Mozilla Data Collective terms; FLEURS is CC-BY-4.0 (Google). The manifests themselves are released CC-BY-4.0.

Citation

bibtex
@misc{antalia1_2026,
  title  = {Antalia 1: An Open Turkish Text-to-Speech Model from a Rights-Clean Pipeline},
  author = {Saygili, Sezgin and Kaplaner, Emre and Ozgul, Oncel and Koktas, Fikri San},
  year   = {2026},
  note   = {Technical report},
  url    = {https://huggingface.co/datasets/cloud0day3/antalia-eval}
}

Contact: sezgin@patientdesk.ai