Machine Translation
medical-bidirectional-machine-translation-checkpoints-511042machine-translationmedical-bidirectional-machine-translation-checkpoints-340694machine_translation_en_vi-GGUFmedical-bidirectional-machine-translationmedical-bidirectional-machine-translation-checkpoints-170348medical-bidirectional-machine-translation-checkpoints-68138Machine_translation_5-10_epochs
NLG-Machine-Translation
SEA Machine Translation
SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.monolingual_machine_translation_datakinyarwanda-english-machine-translation-dataset
Kinyarwanda English Parallel Datasets for Machine translation
A 48,000 Kinyarwanda English Parallel datasets for machine translation, made by curating and translating normal Kinyarwanda sentences into English
shola-machine-translations
SHOLA machine translations
Machine translations of everyday English words into African languages, from
four systems, published so their output can be compared against each other and
against what native speakers actually say.
These are candidate translations, not verified ones. They exist so a
speaker evaluating them at SHOLA has
something to agree with or correct instead of a blank box. Some are wrong. That
is the point: which ones speakers pick is the measurement.… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/shola-machine-translations.igbo_english_machine_translationParallel Igbo-English DatasetMachineTranslation_en_vi
English–Vietnamese Machine Translation Dataset 2025
Dataset Description
This dataset is a large-scale English–Vietnamese parallel corpus designed for training and evaluating neural machine translation systems.
The corpus was constructed by aggregating sentence pairs from several publicly available English–Vietnamese datasets and applying a multi-stage filtering pipeline to reduce malformed samples, duplicated content, language mismatches, and semantically… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/MachineTranslation_en_vi.
