aranemini/central-kurdish-pseudolabel
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.
Central Kurdish → English Pseudo-Labeled Speech Translation Corpus
Dataset Summary
This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).
The dataset was automatically generated using a pipeline composed of:
- Speech segmentation
- Automatic Speech Recognition (ASR)
- Machine Translation (MT)
The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very limited parallel speech translation resources. The corpus was created from publicly available Kurdish speech recordings and automatically translated into English.
Dataset Statistics
The pseudo-labeling pipeline generated approximately 3,200 hours of Kurdish speech aligned with English translations, making it one of the largest speech translation resources available for Kurdish.
Data Quality
The pseudo-labels were generated using state-of-the-art Kurdish speech and language technologies.
The Automatic Speech Recognition (ASR) system employed in the pipeline achieved:
These results represented state-of-the-art performance for Central Kurdish ASR at the time the corpus was created and significantly reduced transcription noise during pseudo-label generation.
Results Obtained Using This Corpus
The pseudo-labeled corpus was subsequently used to train lightweight speech processing models for Central Kurdish.
Automatic Speech Recognition
A lightweight ASR model trained using pseudo-labeled data achieved:
Despite being substantially smaller than multilingual foundation models, the resulting model achieved competitive performance while offering significantly lower computational requirements.
Speech Translation
An end-to-end Speech-to-Text Translation model trained using the pseudo-labeled corpus achieved the following performance:
These results demonstrate that large-scale pseudo-labeling can effectively compensate for the scarcity of manually annotated speech translation corpora in low-resource languages.
Intended Uses
This dataset can be used for:
- End-to-End Speech Translation (S2TT)
- Semi-supervised Speech Translation
- Speech representation learning
- Low-resource speech processing
- Pretraining multilingual speech models
- Research on pseudo-labeling strategies
- Data augmentation for speech translation systems
Limitations
The labels are automatically generated and may contain:
- ASR transcription errors
- Machine translation errors
- Segmentation inaccuracies
Users are encouraged to validate data quality for their specific use cases and, when possible, combine this corpus with manually annotated resources.
Citation
If you use this dataset, please cite:
@inproceedings{mohammadamini25_interspeech,
title = {Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)},
author = {Mohammad Mohammadamini and Aghilas Sini and Marie Tahon and Antoine Laurent},
booktitle = {Interspeech 2025},
year = {2025},
pages = {898--902},
doi = {10.21437/Interspeech.2025-887}
}@inproceedings{mohammadamini2026iwslt,
title = {LIUM Submission for IWSLT 2026 Low-Resource Speech Translation Track},
author = {Mohammad Mohammadamini and Marie Tahon},
year = {2026},
howpublished = {Proceedings of the International Conference on Spoken Language Translation (IWSLT) 2026},
}@inproceedings{mohammadamini2025frugal,
title={Apprentissage de modèles frugaux pour les langues peu dotées à partir de larges modèles d'ASR},
author={Mohammadamini, Mohammad},
school={Le Mans Université},
year={2025}
}