Team Ai
Datasetpublic

BeitTigreAI/tigre-data-parallel-multilingual

Tigre Parallel Corpus This corpus has 66,144 sentence and phrase pairs. Every pair has Tigre (ትግሬ, Ethi_tig) as the source. The targets are Tigrinya, English, Arabic, Swedish and German. The pairs were gathered from three sources: SMOL, a community-contributed file and Tatoeba. Each source was cleaned and put into the same five-column format. Tigre is a Semitic language spoken mainly in Eritrea and written in Ge'ez (Ethiopic) script. Very little parallel data exists for it.… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.

sourceHugging Faceotherupdated 3d agoView on Hugging Face
1likes106downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
BeitTigreAI/tigre-data-parallel-multilingual · Team Ai