Team Ai
Datasetpublic

aranemini/central-kurdish-pseudolabel

Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.

sourceHugging Faceunknownupdated 4mo agoView on Hugging Face
2likes3kdownloads
Dataset Card

Central Kurdish → English Pseudo-Labeled Speech Translation Corpus

Dataset Summary

This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish).

The dataset was automatically generated using a pipeline composed of:

  1. 1.Speech segmentation
  2. 2.Automatic Speech Recognition (ASR)
  3. 3.Machine Translation (MT)

The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very limited parallel speech translation resources. The corpus was created from publicly available Kurdish speech recordings and automatically translated into English.

Dataset Statistics

PropertyValue
Source languageCentral Kurdish (ckb)
Target languageEnglish (en)
Hours of speech~3,200 h
Number of samples~1.7 million
Label typePseudo-labeled

The pseudo-labeling pipeline generated approximately 3,200 hours of Kurdish speech aligned with English translations, making it one of the largest speech translation resources available for Kurdish.

Data Quality

The pseudo-labels were generated using state-of-the-art Kurdish speech and language technologies.

The Automatic Speech Recognition (ASR) system employed in the pipeline achieved:

Evaluation SetWER (%)
Asosoft Test8.18
FLEURS Test20.31

These results represented state-of-the-art performance for Central Kurdish ASR at the time the corpus was created and significantly reduced transcription noise during pseudo-label generation.

Results Obtained Using This Corpus

The pseudo-labeled corpus was subsequently used to train lightweight speech processing models for Central Kurdish.

Automatic Speech Recognition

A lightweight ASR model trained using pseudo-labeled data achieved:

Evaluation SetWER (%)
Asosoft Test7.57
FLEURS Test22.83

Despite being substantially smaller than multilingual foundation models, the resulting model achieved competitive performance while offering significantly lower computational requirements.

Speech Translation

An end-to-end Speech-to-Text Translation model trained using the pseudo-labeled corpus achieved the following performance:

Evaluation SetBLEUChrF++
FLEURS Test21.9752.52
Asosoft Test27.2356.53
COMMUTE Dev25.7351.81
COMMUTE Test21.0949.26

These results demonstrate that large-scale pseudo-labeling can effectively compensate for the scarcity of manually annotated speech translation corpora in low-resource languages.

Intended Uses

This dataset can be used for:

  • —End-to-End Speech Translation (S2TT)
  • —Semi-supervised Speech Translation
  • —Speech representation learning
  • —Low-resource speech processing
  • —Pretraining multilingual speech models
  • —Research on pseudo-labeling strategies
  • —Data augmentation for speech translation systems

Limitations

The labels are automatically generated and may contain:

  • —ASR transcription errors
  • —Machine translation errors
  • —Segmentation inaccuracies

Users are encouraged to validate data quality for their specific use cases and, when possible, combine this corpus with manually annotated resources.

Citation

If you use this dataset, please cite:

bibtex
@inproceedings{mohammadamini25_interspeech,
  title     = {Scaling pseudo-labeling data for end-to-end low-resource speech translation (the case of Kurdish language)},
  author    = {Mohammad Mohammadamini and Aghilas Sini and Marie Tahon and Antoine Laurent},
  booktitle = {Interspeech 2025},
  year      = {2025},
  pages     = {898--902},
  doi       = {10.21437/Interspeech.2025-887}
}
bibtex
@inproceedings{mohammadamini2026iwslt,
  title        = {LIUM Submission for IWSLT 2026 Low-Resource Speech Translation Track},
  author       = {Mohammad Mohammadamini and Marie Tahon},
  year         = {2026},
  howpublished = {Proceedings of the International Conference on Spoken Language Translation (IWSLT) 2026},
}
bibtex
@inproceedings{mohammadamini2025frugal,
  title={Apprentissage de modèles frugaux pour les langues peu dotées à partir de larges modèles d'ASR},
  author={Mohammadamini, Mohammad},
  school={Le Mans Université},
  year={2025}
}