tiny-aya-translate/lahaja-eval
LAHAJA Hindi ASR Eval 3,076 Hindi test utterances (~712 MB) carrying rich speaker metadata — native_language, native_state, gender, age_group, scenario — plus both verbatim and normalized transcripts. Schema in the YAML header above. Held as an evaluation set only: never trained on. Its dialect and native-state labels make it useful for checking whether Hindi ASR quality holds across accents rather than only on the average. This is the benchmark behind hindi-tts-probe, which… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/lahaja-eval.
LAHAJA Hindi ASR Eval
3,076 Hindi test utterances (~712 MB) carrying rich speaker metadata — native_language, native_state, gender, age_group, scenario — plus both verbatim and normalized transcripts. Schema in the YAML header above.
Held as an evaluation set only: never trained on. Its dialect and native-state labels make it useful for checking whether Hindi ASR quality holds across accents rather than only on the average.
This is the benchmark behind [`hindi-tts-probe`](https://github.com/tiny-aya-simultaneous-translation/hindi-tts-probe), which pairs 5 Hindi TTS models with 3 ASR models over 100 samples from this set and measures round-trip WER/CER — the study that selected vasista22/whisper-hindi-large-v2 as the Hindi ASR judge used in the v0.3 evaluation.
Derived from the upstream LAHAJA release; its licence governs redistribution.
Code
Project
TinyAya Stage 2 — Turkish⇄Hindi speech-to-speech translation with a text inner-monologue: a LoRA-adapted Cohere2 backbone driving a frozen Moshi depth decoder over Mimi codes.
The v0.3 run covered 76,250 steps / 2.07 epochs on a Cloud TPU v6e-16 (best val composite 2.8199 @ step 76,000). Read honestly: the text inner-monologue learns to translate (free-run chrF++ ~25.7 / 25.1), while intelligible audio synthesis remains the frontier (ASR-chrF++ 3.7 / 9.6 against a 92.1 / 86.6 ground-truth-audio ceiling) — bounded by the frozen depth decoder, not by translation understanding.
- Results: v0.3 evaluation report
- Training run: W&B `xzcb60bl` · emergence report
- Blog: Adapting Moshi for Low-Resource Speech Translation
Compute for the v0.3 run was provided by Google's TPU Research Cloud (TRC).
