Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.6. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K1 likes255 downloads2d agoHugging Face02bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes115 downloads25d agoHugging Face03ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes108 downloads11mo agoHugging Face04Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes107 downloads3y agoHugging Face05Oumar199 /French_Wolof_Various_Parallel_Corpustexttranslation1K<n<10K3 likes85 downloads2y agoHugging Face06kalixlouiis /HFcourse-english-burmese-parallel-corpus HFcourse-English-Burmese-Parallel-Corpus Dataset Description Dataset Summary The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.texttranslation1K<n<10K12 likes74 downloads6mo agoHugging Face07bekan /english_karakalpak_parallel_corpus_v10 English-Karakalpak Parallel Corpus v10.0 Dataset Description English-Karakalpak Parallel Corpus v10.0 is a high-quality, finalized parallel dataset containing over 50,667 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v10.texttranslation10K<n<100K0 likes74 downloads25d agoHugging Face08PrinceAlhassanNasamu /kusaal-english-parallel-corpus Kusaal-English Parallel Corpus The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset. This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.texttranslation10K<n<100K2 likes62 downloads3mo agoHugging Face09projecte-aina /CA-EN_Parallel_Corpus Dataset Card for CA-EN Parallel Corpus Dataset Description Dataset Summary The CA-EN Parallel Corpus is a Catalan-English dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between English and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EN_Parallel_Corpus.tabulartranslation10M<n<100M1 likes57 downloads1y agoHugging Face10HackHedron /English_Telugu_Parallel_Corpustexttranslation100K<n<1M1 likes54 downloads1y agoHugging Face11abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes50 downloads4mo agoHugging Face12bekan /english_karakalpak_parallel_corpus_v8 English-Karakalpak Parallel Corpus v8.0 Dataset Description English-Karakalpak Parallel Corpus v8.0 is a high-quality, finalized parallel dataset containing over 37,257 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v8.texttranslation10K<n<100K1 likes50 downloads3mo agoHugging Face13DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes49 downloads4mo agoHugging Face14mbaye930 /wolof-arabic-parallel-corpus MudawanSn: A Gold-Standard Wolof--Arabic Parallel Corpus for Machine Translation A publicly available parallel corpus for the Wolof–Arabic language pair, a gold-standard resource containing 1,271 sentence-aligned pairs. The corpus consists of manual translations from Wolof into Modern Standard Arabic (MSA). The source texts are drawn from the MasakhaNER corpus, covering politics, society, religion, and sports in Senegalese news discourse. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/mbaye930/wolof-arabic-parallel-corpus.texttranslation1K<n<10K3 likes39 downloads4mo agoHugging Face15bekan /english_karakalpak_parallel_corpus_v7 English-Karakalpak Parallel Corpus v7.0 Dataset Description English-Karakalpak Parallel Corpus v7.0 is a high-quality, finalized parallel dataset containing over 32,972 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v7.texttranslation10K<n<100K0 likes38 downloads6mo agoHugging Face16rakhine-nlp /rakhine-english-parallel-corpus 🌐 Rakhine–English Parallel Corpus A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP). 🎯 Purpose This dataset is designed to support: Machine Translation (MT) Neural Machine Translation (NMT) Language Modeling Low-resource NLP research Linguistic and dialect studies Language preservation and documentation 📌 Overview Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.texttranslationn<1K0 likes36 downloads3mo agoHugging Face17geekdiop /A-Wolof-Arabic-Parallel-Corpustext1K<n<10K1 likes30 downloads24d agoHugging Face18bekan /english_karakalpak_parallel_corpus_v6 English-Karakalpak Parallel Corpus v6.0 Dataset Description English-Karakalpak Parallel Corpus v6.0 is a high-quality, finalized parallel dataset containing over 31,000 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v6.texttranslation10K<n<100K0 likes29 downloads6mo agoHugging Face19navinaananthan /Kurdish-Sorani-Parallel-Corpustext100K<n<1M5 likes27 downloads3y agoHugging Face20bekan /english_karakalpak_parallel_corpus_v9 English-Karakalpak Parallel Corpus v9.0 Dataset Description English-Karakalpak Parallel Corpus v9.0 is a high-quality, finalized parallel dataset containing over 40,513 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v9.texttranslation10K<n<100K0 likes24 downloads3mo agoHugging Face21Bapynshngain /English-Khasi-Parallel-Corpus-v1gateddata_source: Web-scraped data manually vetted by me. NIT Silchar’s "EnKhCorp1.0: An English–Khasi Corpus." Tatoeba project. samanantar : the largest publicly available parallel corpora collection for 11 indic languages acknowledgments: Special thanks to Ahlad, NIT Silchar for their "EnKhCorp1.0" dataset, IIT Madras and other institutions involved in the creation of Samanantar and to the contributors of the Tatoeba project. texttranslation10K<n<100K0 likes22 downloads1y agoHugging Face22VIITPune /Deshika-Maharashtri_Prakrit_to_English_Parallel_CorpusMaharashtri Prakrit to English Parallel Corpus Dataset Summary This dataset contains parallel text data for translating from Maharashtri Prakrit (an ancient Indo-Aryan language) to English. It is designed to aid in developing machine translation systems, language models, and linguistic research for this underrepresented language. The dataset is collected from historical texts, scriptures, and scholarly resources. Key Features: Source Language: Maharashtri Prakrit Target Language: English… See the full description on the dataset page: https://huggingface.co/datasets/VIITPune/Deshika-Maharashtri_Prakrit_to_English_Parallel_Corpus.texttranslation1K<n<10K1 likes20 downloads2y agoHugging Face23tachiwin /multilingual_parallel_corpustext10K<n<100K0 likes20 downloads26d agoHugging Face24fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes19 downloads1mo agoHugging Face25Omarrran /kashmiri_English_parallel_corpus_49Kgated license: apache-2.0 task_categories: translation language: ks Usage Terms for this Dataset Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications. Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following: @misc {haq_nawaz_malik_2024, author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.text10K<n<100K2 likes17 downloads1y agoHugging Face26TankuVie /ted_talks_multilingual_parallel_corpustext10K<n<100K2 likes16 downloads3y agoHugging Face27navinaananthan /Dhivehi-English-ParallelCorpustext100K<n<1M1 likes16 downloads3y agoHugging Face28Bapynshngain /Bapyn-EnKha-Parallel-CorpusgatedThis dataset is archived on Zenodo: DOI: https://doi.org/10.5281/zenodo.18703428 Note: This dataset will be readily granted access upon request. The form is only to log users and make sure they adhere to the license terms and conditions To request access: Submit an access request via the Hugging Face dataset page. Complete the following access form: 👉 Request Access Form Please note that requests will be reviewed and verified by matching your Hugging Face username and details with the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/Bapyn-EnKha-Parallel-Corpus.texttranslation10K<n<100K1 likes16 downloads8mo agoHugging Face29kalixlouiis /pali-myanmar-parallel-corpus-1ktexttranslation1K<n<10K5 likes15 downloads5mo agoHugging Face30MEDHARVIX-SYSTEMS /bhasaflow-khasi-english-parallel-corpus-v1 BhasaFlow Khasi-English Parallel Corpus v1 By Medharvix Systems Private Limited Overview A curated parallel corpus of Khasi-English sentence pairs designed for machine translation research and development, with a focus on low-resource language technology for Northeast India. Dataset Structure Column Description sentence_id Unique sentence identifier english_text English sentence khasi_text Khasi translation Usage from datasets… See the full description on the dataset page: https://huggingface.co/datasets/MEDHARVIX-SYSTEMS/bhasaflow-khasi-english-parallel-corpus-v1.texttranslationn<1K19 likes14 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.