Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.5. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K1 likes218 downloads4d agoHugging Face02Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes108 downloads3y agoHugging Face03bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes104 downloads23d agoHugging Face04ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes102 downloads11mo agoHugging Face05Oumar199 /French_Wolof_Various_Parallel_Corpustexttranslation1K<n<10K3 likes82 downloads2y agoHugging Face06PrinceAlhassanNasamu /kusaal-english-parallel-corpus Kusaal-English Parallel Corpus The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset. This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.texttranslation10K<n<100K2 likes70 downloads3mo agoHugging Face07bekan /english_karakalpak_parallel_corpus_v10 English-Karakalpak Parallel Corpus v10.0 Dataset Description English-Karakalpak Parallel Corpus v10.0 is a high-quality, finalized parallel dataset containing over 50,667 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v10.texttranslation10K<n<100K0 likes70 downloads23d agoHugging Face08kalixlouiis /HFcourse-english-burmese-parallel-corpus HFcourse-English-Burmese-Parallel-Corpus Dataset Description Dataset Summary The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.texttranslation1K<n<10K12 likes69 downloads6mo agoHugging Face09Omarrran /kashmiri_English_parallel_corpus_49Kgated license: apache-2.0 task_categories: translation language: ks Usage Terms for this Dataset Purpose of UseThis dataset is made available for the purpose of training machine learning models, academic research, and other non-commercial uses and its applications. Citation RequirementIf you use this dataset for research, training models, or any other purpose, you must provide proper attribution by citing the following: @misc {haq_nawaz_malik_2024, author = { {HAQ NAWAZ… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/kashmiri_English_parallel_corpus_49K.text10K<n<100K2 likes63 downloads1y agoHugging Face10projecte-aina /CA-EN_Parallel_Corpus Dataset Card for CA-EN Parallel Corpus Dataset Description Dataset Summary The CA-EN Parallel Corpus is a Catalan-English dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between English and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EN_Parallel_Corpus.tabulartranslation10M<n<100M1 likes54 downloads1y agoHugging Face11DatarrX /Myanmar-Written-Spoken-Parallel-Corpus Myanmar Written-Spoken Parallel Corpus (MWSPC) Dataset Description Myanmar Written-Spoken Parallel Corpus (MWSPC) is a high-quality open-source dataset designed to bridge the gap between formal written Burmese and daily spoken Burmese. This dataset is crucial for building natural-sounding AI models that understand the linguistic nuances of the Myanmar language. Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX | ဒေတာ-အက်စ် Language: Burmese… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Myanmar-Written-Spoken-Parallel-Corpus.texttext-generation1K<n<10K6 likes50 downloads4mo agoHugging Face12bekan /english_karakalpak_parallel_corpus_v8 English-Karakalpak Parallel Corpus v8.0 Dataset Description English-Karakalpak Parallel Corpus v8.0 is a high-quality, finalized parallel dataset containing over 37,257 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v8.texttranslation10K<n<100K1 likes49 downloads3mo agoHugging Face13HackHedron /English_Telugu_Parallel_Corpustexttranslation100K<n<1M1 likes47 downloads1y agoHugging Face14abdelhaqueidali /Kabyle-Latin-to-Tifinagh-Parallel-Corpus Dataset Card for Kabyle Latin-to-Tifinagh Parallel Corpus This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-based orthography with automated, context-aware Amazigh script transliterations. It is built by processing raw text data through a rule-based algorithmic pipeline designed to enforce strict orthographic purity, manage contextual phonetic mutations, and isolate foreign vocabulary. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus.texttranslation1M<n<10M0 likes46 downloads3mo agoHugging Face15rakhine-nlp /rakhine-english-parallel-corpus 🌐 Rakhine–English Parallel Corpus A parallel corpus for Rakhine ↔ English machine translation, low-resource language research, and Natural Language Processing (NLP). 🎯 Purpose This dataset is designed to support: Machine Translation (MT) Neural Machine Translation (NMT) Language Modeling Low-resource NLP research Linguistic and dialect studies Language preservation and documentation 📌 Overview Rakhine is spoken by millions of people in… See the full description on the dataset page: https://huggingface.co/datasets/rakhine-nlp/rakhine-english-parallel-corpus.texttranslationn<1K0 likes36 downloads3mo agoHugging Face16mbaye930 /wolof-arabic-parallel-corpus MudawanSn: A Gold-Standard Wolof--Arabic Parallel Corpus for Machine Translation A publicly available parallel corpus for the Wolof–Arabic language pair, a gold-standard resource containing 1,271 sentence-aligned pairs. The corpus consists of manual translations from Wolof into Modern Standard Arabic (MSA). The source texts are drawn from the MasakhaNER corpus, covering politics, society, religion, and sports in Senegalese news discourse. Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/mbaye930/wolof-arabic-parallel-corpus.texttranslation1K<n<10K3 likes35 downloads4mo agoHugging Face17bekan /english_karakalpak_parallel_corpus_v7 English-Karakalpak Parallel Corpus v7.0 Dataset Description English-Karakalpak Parallel Corpus v7.0 is a high-quality, finalized parallel dataset containing over 32,972 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v7.texttranslation10K<n<100K0 likes33 downloads6mo agoHugging Face18navinaananthan /Kurdish-Sorani-Parallel-Corpustext100K<n<1M5 likes31 downloads3y agoHugging Face19geekdiop /A-Wolof-Arabic-Parallel-Corpustext1K<n<10K1 likes29 downloads22d agoHugging Face20bekan /english_karakalpak_parallel_corpus_v6 English-Karakalpak Parallel Corpus v6.0 Dataset Description English-Karakalpak Parallel Corpus v6.0 is a high-quality, finalized parallel dataset containing over 31,000 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT Script: Latin… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v6.texttranslation10K<n<100K0 likes29 downloads6mo agoHugging Face21bekan /english_karakalpak_parallel_corpus_v9 English-Karakalpak Parallel Corpus v9.0 Dataset Description English-Karakalpak Parallel Corpus v9.0 is a high-quality, finalized parallel dataset containing over 40,513 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v9.texttranslation10K<n<100K0 likes27 downloads3mo agoHugging Face22Bapynshngain /English-Khasi-Parallel-Corpus-v1gateddata_source: Web-scraped data manually vetted by me. NIT Silchar’s "EnKhCorp1.0: An English–Khasi Corpus." Tatoeba project. samanantar : the largest publicly available parallel corpora collection for 11 indic languages acknowledgments: Special thanks to Ahlad, NIT Silchar for their "EnKhCorp1.0" dataset, IIT Madras and other institutions involved in the creation of Samanantar and to the contributors of the Tatoeba project. texttranslation10K<n<100K0 likes20 downloads1y agoHugging Face23tachiwin /multilingual_parallel_corpustext10K<n<100K0 likes19 downloads24d agoHugging Face24fahim-ling /Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLPgated Sylheti-Bangla-English-Russian-German Parallel Corpus for NLP Welcome to the first open-source multilingual parallel corpus for the Sylheti (syl) language, engineered by a native Linguistics student. This dataset bridges Sylheti with four major global high-resource languages spanning three distinct language families (Indo-Aryan, Germanic, and Slavic) to support Computational Linguistics (CL), Natural Language Processing (NLP) research, and Large Language Model (LLM) fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/fahim-ling/Sylheti-Bangla-English-Russian-German-Parallel-Corpus-for-NLP.texttranslationn<1K1 likes19 downloads1mo agoHugging Face25VIITPune /Deshika-Maharashtri_Prakrit_to_English_Parallel_CorpusMaharashtri Prakrit to English Parallel Corpus Dataset Summary This dataset contains parallel text data for translating from Maharashtri Prakrit (an ancient Indo-Aryan language) to English. It is designed to aid in developing machine translation systems, language models, and linguistic research for this underrepresented language. The dataset is collected from historical texts, scriptures, and scholarly resources. Key Features: Source Language: Maharashtri Prakrit Target Language: English… See the full description on the dataset page: https://huggingface.co/datasets/VIITPune/Deshika-Maharashtri_Prakrit_to_English_Parallel_Corpus.texttranslation1K<n<10K1 likes18 downloads2y agoHugging Face26TankuVie /ted_talks_multilingual_parallel_corpustext10K<n<100K2 likes16 downloads3y agoHugging Face27kalixlouiis /pali-myanmar-parallel-corpus-1ktexttranslation1K<n<10K5 likes15 downloads5mo agoHugging Face28navinaananthan /Dhivehi-English-ParallelCorpustext100K<n<1M1 likes14 downloads3y agoHugging Face29Bapynshngain /Bapyn-EnKha-Parallel-CorpusgatedThis dataset is archived on Zenodo: DOI: https://doi.org/10.5281/zenodo.18703428 Note: This dataset will be readily granted access upon request. The form is only to log users and make sure they adhere to the license terms and conditions To request access: Submit an access request via the Hugging Face dataset page. Complete the following access form: 👉 Request Access Form Please note that requests will be reviewed and verified by matching your Hugging Face username and details with the… See the full description on the dataset page: https://huggingface.co/datasets/Bapynshngain/Bapyn-EnKha-Parallel-Corpus.texttranslation10K<n<100K1 likes14 downloads8mo agoHugging Face30bekan /english_karakalpak_parallel_corpus_v3-4 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus v3-4 is a high-quality dataset containing 2,722 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v3-4.texttranslation1K<n<10K0 likes14 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.