BeitTigreAI/tigre-data-parallel-multilingual
Tigre Parallel Corpus This corpus has 66,144 sentence and phrase pairs. Every pair has Tigre (ትግሬ, Ethi_tig) as the source. The targets are Tigrinya, English, Arabic, Swedish and German. The pairs were gathered from three sources: SMOL, a community-contributed file and Tatoeba. Each source was cleaned and put into the same five-column format. Tigre is a Semitic language spoken mainly in Eritrea and written in Ge'ez (Ethiopic) script. Very little parallel data exists for it.… See the full description on the dataset page: https://huggingface.co/datasets/BeitTigreAI/tigre-data-parallel-multilingual.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face