BeitTigreAI/tigre-data-parallel-multilingual
Tigre Parallel Corpus This corpus has 66,144 sentence and phrase pairs. Every pair has Tigre (ትግሬ, Ethi_tig) as the source. The targets are Tigrinya, English, Arabic, Swedish and German. The pairs were gathered from three sources: SMOL, a community-contributed file and Tatoeba. Each source was cleaned and put into the same five-column format. Tigre is a Semitic language spoken mainly in Eritrea and written in Ge'ez (Ethiopic) script. Very little parallel data exists for it.… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.
Tigre Parallel Corpus
This corpus has 66,144 sentence and phrase pairs. Every pair has Tigre (ትግሬ, `Ethi_tig`) as the source. The targets are Tigrinya, English, Arabic, Swedish and German. The pairs were gathered from three sources: SMOL, a community-contributed file and Tatoeba. Each source was cleaned and put into the same five-column format.
Tigre is a Semitic language spoken mainly in Eritrea and written in Ge'ez (Ethiopic) script. Very little parallel data exists for it.
Contents
Pairs by target language (`all`): Arabic 27,866 · English 19,901 · Swedish 9,003 · Tigrinya 7,741 · German 1,633
Format
All files are UTF-8 CSV with the same five columns:
Language codes use the form Script_language (ISO 15924 script + ISO 639-3 language).
Usage
from datasets import load_dataset
ds = load_dataset("BeitTigreAI/tigre-data-parallel-multilingual") # everything
ds = load_dataset("BeitTigreAI/tigre-data-parallel-multilingual", "smol_tig_en") # one source
# filter by target language or source
en = ds["train"].filter(lambda x: x["tgt_lang"] == "Latn_eng")
pro = ds["train"].filter(lambda x: x["source"] == "smol")There is a single train split. Because some Tigre sentences appear several times (see Limitations), any evaluation set you carve out should be split by Tigre sentence. That way the same sentence never appears in both training and test data.
How the data was built and cleaned
SMOL (smol_tig_ti, smol_tig_en)
These come from Google's SMOL dataset: the SmolSent, SmolDoc and GATITOS subsets.
- Tigre → English: the English→Tigre files were reversed so that Tigre is the source. SmolDoc documents were split into sentence pairs.
- Tigre → Tigrinya: SMOL has no direct Tigre–Tigrinya data. It does have English→Tigre and English→Tigrinya translations of the same English text, so the Tigre and Tigrinya translations of each English item were paired:
- SmolSent sentences were joined on their sentence ID.
- SmolDoc documents were joined on document ID and aligned sentence by sentence. 3 documents with mismatched sentence counts were excluded.
- GATITOS entries have no IDs, so they were joined on an exact match of the English entry.
- Cleanup (both files): for GATITOS entries with several Tigre translations, only the first was kept. Whitespace was trimmed, and empty and exact-duplicate pairs were removed.
Community (community_tig_en)
Volunteer English→Tigre translations, originally 3,006 lines in English ||| Tigre format. The file covers two domains: software and wiki interface strings (315 pairs) and Eritrean history lesson material (1,990 pairs). 701 lines were removed:
- Removed:
- 251 single-word entries
- 213 exact duplicates
- 168 interface strings that are templates rather than sentences (
$1,{{PLURAL:…}},[[links]], HTML) - 38 history rows with OCR noise (stray Latin letters inside the Tigre text)
- 22 history rows whose lengths suggest a partial translation or misalignment
- 8 interface messages that were cut off and misaligned
- 1 untranslated row
- Fixed:
- Spreadsheet-style quote marks and extra spaces were removed, and text was normalized to Unicode NFC.
- 23 run-together English words were split after checking each one, e.g. thedevelopmental → the developmental, skillstraining → skills training.
Tatoeba (tatoeba_tig_multi)
Tatoeba's "Sentence pairs" exports for Tigre–Arabic, Tigre–English, Tigre–Swedish and Tigre–German, downloaded 2026-10-03 (52,065 pairs). Each language pair was cleaned separately with the same rules, and the results were combined. 3,961 pairs were removed:
- Removed:
- 3,722 single-word entries
- 178 length mismatches, measured against the median length ratio for each language. Most join several alternative Tigre versions of one short sentence.
- 27 exact duplicates
- 19 rows in the wrong script
- 12 rows with English words or stray Latin letters inside the Tigre
- 3 untranslated rows
- Fixed:
- "::" typed in place of the Ethiopic full stop was replaced with "።" (9,967 sentences).
- Spaces before Ethiopic punctuation were removed.
- Words in parentheses inside Tigre sentences, such as synonyms, glosses or pronoun hints, were removed when the target had none (184 sentences).
Full per-source details are in `docs/`.
Limitations
- Mixed quality. SMOL is professionally translated. The community and Tatoeba data are volunteer and crowd-sourced, and translation quality varies. Consider SMOL for evaluation.
- The Tigre–Tigrinya pairs are indirect. Each side was translated from English separately, so the two sides can differ more than direct translations would.
- Tigrinya-style spelling. Some community and Tatoeba Tigre text uses Tigrinya-style spelling or vocabulary (e.g. ሓ/ኣ/ዓ in place of ሐ/አ/ዐ). After review, this was kept as valid Tigre usage.
- Short, conversational text. Tatoeba sentences are short (typically about 4 words). GATITOS entries are words and short phrases without context.
- Translated Tigre. In SMOL and much of Tatoeba, the Tigre was translated from another language rather than written originally in Tigre.
- Community data artifacts. The history text appears to be OCR-derived and may still contain typos. Its English is mostly lowercase.
- Repeated Tigre sentences. The same Tigre sentence can appear more than once: with different targets, with several Tatoeba translations, or in SMOL Tigre→English and Tigre→Tigrinya. There are no exact duplicate rows.
- Arabic code. Arabic is coded
Arab_ara(macrolanguage), as Tatoeba labels it. Most of it is Modern Standard Arabic (Arab_arb), but some is dialectal.
Licenses
Each part keeps the license of its source:
Attribution to SMOL (Google) and to Tatoeba and its contributors is required by their licenses. Credit to the community contributors is appreciated but not required.
Citation
If you use the SMOL portions, please cite:
@misc{caswell2025smol,
title = {SMOL: Professionally translated parallel data for 115 under-represented languages},
author = {Caswell, Isaac and others},
year = {2025},
eprint = {2502.12301},
archivePrefix = {arXiv}
}
@inproceedings{jones2023gatitos,
title = {GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation},
author = {Jones, Alex and others},
booktitle = {Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing},
year = {2023},
url = {https://aclanthology.org/2023.emnlp-main.26/}
}Tatoeba: <https://tatoeba.org>
