Team Ai
20 results

Machine Translation

aisingapore /NLG-Machine-Translationgated SEA Machine Translation SEA Machine Translation evaluates a model's ability to translate a document from a source language into a target language coherently and fluently. It is sampled from FLORES 200 for Burmese, Chinese, English, Indonesian, Khmer, Malay, Tamil, Thai, and Vietnamese, and NusaX for Indonesian, Javanese, and Sundanese. Supported Tasks and Leaderboards SEA Machine Translation is designed for evaluating chat or instruction-tuned large language models… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Machine-Translation.texttext-generation10K<n<100K7 likes1.4k downloads9mo agoHugging FaceDigitalUmuganda /monolingual_machine_translation_datatext100K<n<1M0 likes287 downloads3y agoHugging FaceDigitalUmuganda /kinyarwanda-english-machine-translation-dataset Kinyarwanda English Parallel Datasets for Machine translation A 48,000 Kinyarwanda English Parallel datasets for machine translation, made by curating and translating normal Kinyarwanda sentences into English 5 likes204 downloads4y agoHugging FaceAfriSpeech /shola-machine-translations SHOLA machine translations Machine translations of everyday English words into African languages, from four systems, published so their output can be compared against each other and against what native speakers actually say. These are candidate translations, not verified ones. They exist so a speaker evaluating them at SHOLA has something to agree with or correct instead of a blank box. Some are wrong. That is the point: which ones speakers pick is the measurement.… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/shola-machine-translations.texttranslation10M<n<100M0 likes199 downloads15d agoHugging Faceignatius /igbo_english_machine_translationParallel Igbo-English Datasettranslation10K<n<100K6 likes183 downloads3y agoHugging FaceTran1312 /MachineTranslation_en_vi English–Vietnamese Machine Translation Dataset 2025 Dataset Description This dataset is a large-scale English–Vietnamese parallel corpus designed for training and evaluating neural machine translation systems. The corpus was constructed by aggregating sentence pairs from several publicly available English–Vietnamese datasets and applying a multi-stage filtering pipeline to reduce malformed samples, duplicated content, language mismatches, and semantically… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/MachineTranslation_en_vi.texttranslation10M<n<100M5 likes131 downloads15d agoHugging Face