Team Ai
Modelpublic

SwarajSolanke-turtle/Marathi_Text_To_Speech

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes164downloads
Model Card

IndicF5 Marathi – Numbers v2

IndicF5-Marathi-Numbers-v2 is a fine-tuned version of ai4bharat/IndicF5 for Marathi text-to-speech, focused on reading numbers correctly: amounts (९८,००० रुपये), lakhs and crores (73 लाखांच्या), times (७:२०), dates and percentages. It does zero-shot voice cloning from a short reference clip, like the base model.

It was built for customer-facing Marathi speech in insurance and finance, such as premium quotes, policy details and callback times, where base models often skip or misread numbers.

Developed byTurtlemint
Model typeNon-autoregressive flow-matching TTS (F5-TTS Base, Diffusion Transformer)
LanguageMarathi (mr)
Fine-tuned fromai4bharat/IndicF5
Checkpointmodel_48000.pt (48,000 updates)
Output24 kHz mono audio (Vocos vocoder)

Highlights

  • —🔢 Reads numbers in Marathi: trained on 5,000 extra sentences with Devanagari and Latin digits, comma-grouped amounts, times and dates in context.
  • —🗣️ Zero-shot voice cloning: give a 5–10 s reference clip and its transcript.
  • —🇮🇳 Natural Marathi prosody: keeps the base IndicF5 Marathi quality, further tuned on about 31.5 h of Marathi speech.
  • —⚡ Fast, non-autoregressive inference: no autoregressive decoding, so speed stays stable for long sentences.

Quick start

Installation

bash
pip install git+https://github.com/ai4bharat/IndicF5.git
pip install huggingface_hub soundfile

Inference

python
import soundfile as sf
from huggingface_hub import hf_hub_download
from f5_tts.model import DiT
from f5_tts.infer.utils_infer import (
    load_model, load_checkpoint, load_vocoder,
    preprocess_ref_audio_text, infer_process,
)

REPO_ID = "SwarajSolanke-turtle/Marathi_Text_To_Speech"
ckpt  = hf_hub_download(REPO_ID, "model_48000.pt")
vocab = hf_hub_download(REPO_ID, "vocab.txt")

device = "cuda"
model_cfg = dict(dim=1024, depth=22, heads=16, ff_mult=2, text_dim=512, conv_layers=4)

model = load_model(DiT, model_cfg, mel_spec_type="vocos", vocab_file=vocab, device=device)
model = load_checkpoint(model, ckpt, device, use_ema=True)
vocoder = load_vocoder(vocoder_name="vocos", device=device)

ref_audio, ref_text = preprocess_ref_audio_text(
    "reference.wav",                       # 5–10 s clean Marathi speech
    "संदर्भ ऑडिओमध्ये बोललेला मजकूर येथे लिहा.",  # its exact transcript
)

text = "७३ लाखांच्या कव्हरसाठी वार्षिक प्रीमियम सुमारे ९८,००० रुपये येतो."
wav, sr, _ = infer_process(ref_audio, ref_text, text, model, vocoder,
                           mel_spec_type="vocos", device=device)

sf.write("output.wav", wav, sr)
Tip: results are best when the reference clip is clean (no music or noise), 5–10 s long, and its transcript is exact.

Training details

Training data

DatasetUtterancesDurationDescription
Marathi speech (original)10,93922.37 hClean Marathi read speech
Numbers v12,7913.24 hSynthetic sentences with numbers
Numbers v2 (new)5,0005.95 hWider set of number formats: amounts, lakhs/crores, times, dates, mixed scripts
Total (merged)18,73031.56 h

Numbers v2 clips are 1.7–8.1 s long (mean 4.3 s). The sentences are templated, domain-specific Marathi text about insurance, payments and scheduling. Audio was generated with an earlier Marathi fine-tune of this model and mixed with real speech so the model doesn't drift.

Training procedure

HyperparameterValue
ArchitectureDiT: dim 1024, depth 22, 16 heads, FF mult 2, text dim 512, 4 conv layers
ObjectiveConditional flow matching (masked mel infilling)
TokenizerCustom IndicF5 character vocabulary
Effective batch size32 samples (2 per GPU × 16 grad accumulation)
OptimizerAdamW
Learning rate1e-5 with linear warmup and decay (≈9.65e-6 → 6.73e-6 over this stage)
Total updates48,000
EMAYes (use EMA weights at inference)
Hardware1 × NVIDIA A10G (24 GB)
Training time (final stage)~2.8 h for updates 8k → 48k

Evaluation

Training loss

Flow-matching loss on the training set, averaged over 4,000-update windows. Each step samples a random noise level, so single-step loss is noisy; the window averages are the meaningful numbers.

UpdatesMean lossMin loss
8k – 12k0.56360.2397
12k – 16k0.56890.2238
16k – 20k0.57030.2393
20k – 24k0.56520.2313
24k – 28k0.55930.2317
28k – 32k0.55030.2186
32k – 36k0.56030.2247
36k – 40k0.56250.2267
40k – 44k0.55080.2236
44k – 48k0.55580.2171
Summary at 48k updatesValue
Mean loss (last 1k updates)0.556 ± 0.377
Best single-step loss0.2171
Final learning rate6.73e-6

The loss levels off at about 0.55–0.56, so the model has converged at this checkpoint.

Objective metrics

Held-out intelligibility (CER/WER from a Marathi ASR model) and naturalness (MOS) results are not published yet. They will be added to this card when available.


Intended use

Intended for

  • —Marathi voice assistants, IVR systems and customer-communication audio
  • —Reading numbers aloud in Marathi: prices, premiums, policy numbers, dates, times
  • —Research on Indic TTS and text normalization

Out of scope

  • —Imitating real people without their explicit consent
  • —Fraud, impersonation, disinformation, or any audio presented as a real person's speech without disclosure
  • —Languages other than Marathi (the base model supports more Indic languages, but this fine-tune targets Marathi only)

Limitations and bias

  • —Synthetic number data: part of the training audio was produced by an earlier version of this model, so its artifacts and pronunciation habits may carry over.
  • —Domain bias: number sentences are mostly about insurance and finance, so general or literary text may sound less natural.
  • —Very long numbers: IDs, phone numbers and long digit strings are best expanded or spaced out before synthesis.
  • —Depends on the reference clip: a noisy or wrongly transcribed reference lowers quality and can cause skipped words.
  • —Code-mixed text: Marathi mixed with English words is only partly supported.

Ethical considerations

This model can clone voices from a few seconds of audio. Use it only with reference audio you have permission to use, and disclose clearly when audio is AI-generated. Do not use it to impersonate anyone or to deceive.


Citation

If you use this model, please cite the original F5-TTS and IndicF5 works:

bibtex
@article{chen2024f5tts,
  title   = {F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching},
  author  = {Chen, Yushen and Niu, Zhikang and Ma, Ziyang and Deng, Keqi and Wang, Chunhui and Zhao, Jian and Yu, Kai and Chen, Xie},
  journal = {arXiv preprint arXiv:2410.06885},
  year    = {2024}
}

@misc{AI4Bharat_IndicF5_2025,
  author       = {Praveen S V and Srija Anand and Soma Siddhartha and Mitesh M. Khapra},
  title        = {IndicF5: High-Quality Text-to-Speech for Indian Languages},
  year         = {2025},
  url          = {https://github.com/AI4Bharat/IndicF5}
}

Acknowledgements

Note

  • —model do not understand the complex number or the number on which model is not trained if this is case then performed the normalization of digit into words and then send to the model , it will then work well on the data you have