ai-sage/GigaAM-Multilingual
GigaAM Multilingual
GigaAM Multilingual is a family of Conformer-based foundation models (220M / 600M parameters) pre-trained with a HuBERT-style objective on 2M hours of speech across 70+ languages and fine-tuned for speech recognition with character-wise CTC decoders on 50K hours.
The models provide best-in-class open-source quality on Russian, Kazakh, Kyrgyz, and Uzbek, and moderate quality on English.
GigaAM Multilingual includes the following model variants:
ssl— 220M self-supervised encoderctc— 220M ASR model with a character-wise CTC decoderlarge_ssl— 600M self-supervised encoderlarge_ctc— 600M ASR model with a character-wise CTC decoder
Model Performance
Word Error Rate (%) on Common Voice (CV), FLEURS, and internal in-the-wild test sets. Utterances longer than 30 s and references containing digits are excluded; references/hypotheses are normalized (lowercasing, punctuation removal, numerals→words); greedy decoding. Best per row in bold.
Usage
from transformers import AutoModel
revision = "ctc" # any variant: ssl, ctc, large_ssl, large_ctc
model = AutoModel.from_pretrained(
"ai-sage/GigaAM-Multilingual",
revision=revision,
trust_remote_code=True,
)
transcription = model.transcribe("example.wav")
print(transcription)Recommended versions:
torch==2.10.*,torchaudio==2.10.*transformers==5.*- (any)
hydra-core,omegaconf
Full usage guide can be found in the example.
Fine-tuning to a new language
The ssl / large_ssl backbones can be adapted to a new language — see the fine-tuning guide and the example notebook.
Citation
@misc{gigaam_multilingual,
title={GigaAM Multilingual: Foundation Model for Underrepresented Languages},
author={Andrei Kuzmenko and Alexandr Maximenko and Aleksandr Kutsakov and Georgii Gospodinov and Dmitrii Bolotov and Oleg Kutuzov and Pavel Bogomolov and Fyodor Minkin},
year={2026},
eprint={2607.10371},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2607.10371}
}