Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cloverx-id /lumi-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/lumi-repository-parallel-en-id-corpus.tabulartranslation100M<n<1B2 likes2.5k downloads6h agoHugging Face02browndw /human-ai-parallel-corpus Human-AI Parallel English Corpus (HAP-E) 🙃 Purpose The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs). The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words. Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.texttext-classification10K<n<100K5 likes1.7k downloads2y agoHugging Face03Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes302 downloads18d agoHugging Face04mteb /english-danish-parallel-corpus DanishMedicinesAgencyBitextMining An MTEB dataset Massive Text Embedding Benchmark A Bilingual English-Danish parallel corpus from The Danish Medicines Agency. Task category t2t Domains Medical, Written Reference https://sprogteknologi.dk/dataset/bilingual-english-danish-parallel-corpus-from-the-danish-medicines-agency How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/english-danish-parallel-corpus.texttranslation10K<n<100K0 likes248 downloads1y agoHugging Face05Okwu /african-language-parallel-corpus African Language Parallel Corpus Human-created, human-validated parallel sentence pairs for three African languages, released openly by Okwu. Version 1.6. Dataset summary A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own language-learning curriculum — content authored and reviewed by native-speaker educators — supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.texttranslation10K<n<100K1 likes218 downloads2h agoHugging Face06raptorkwok /cantonese-chinese-parallel-corpus-baseThis is a dataset of Cantonese-Written Chinese Parallel Corpus, containing 130k+ pairs of Cantonese and Traditional Chinese parallel sentences. texttranslation100K<n<1M17 likes205 downloads3y agoHugging Face07browndw /human-ai-parallel-corpus-biber Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.tabular10K<n<100K0 likes200 downloads2y agoHugging Face08browndw /human-ai-parallel-corpus-2 Human-AI Parallel English Corpus-2 (HAP-E-2) 🙃 Purpose The HAP-E-2 corpus is an extension of the original HAP-E corpus, with the addition of updated models. is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs). The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words. Thus, a second 500-word chunk… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-2.texttext-classification100K<n<1M0 likes174 downloads8mo agoHugging Face09adiga-ai /circassian-parallel-corpus Circassian-Russian Parallel Corpus v1.0 This is a high-quality dataset containing over 330,000 parallel text pairs for machine translation between Russian and the Circassian language in its two literary dialects: East Circassian (Kabardian, kbd) and West Circassian (Adyghe, ady). About Circassian Circassian is an indigenous language of the Northwest Caucasus region. The language is notable for its complex phonological system (featuring 50+ consonants)… See the full description on the dataset page: https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus.texttranslation100K<n<1M5 likes165 downloads1y agoHugging Face10Moleys /Filtered-Japanese-English-Parallel-Corpusdef prompt(japanese, english): system_prompt = cleandoc("""<s>[INST]Your role is to evaluate the accuracy of the provided Japanese to English translation. - Translations with parts missing should be rejected. - Incomplete translations should be rejected. - Inaccurate translations should be rejected. - Poor grammar should be rejected. - Any kind of mistake should be rejected. - Bad spelling should be rejected. - Low quality english should be rejected. - Low… See the full description on the dataset page: https://huggingface.co/datasets/Moleys/Filtered-Japanese-English-Parallel-Corpus.texttranslation10M<n<100M5 likes137 downloads2y agoHugging Face11projecte-aina /ES-AN_Parallel_Corpus Dataset Card for ES-AN Parallel Corpus Dataset Summary The ES-AN Parallel Corpus is a Spanish-Aragonese dataset created to support the use of under-resourced languages from Spain, such as Aragonese, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Aragonese and Spanish in any direction, as well as Multilingual Machine Translation models.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-AN_Parallel_Corpus.texttranslation10K<n<100K2 likes128 downloads1y agoHugging Face12browndw /human-ai-parallel-corpus-docuscope COCA-AI Parallel Corpus (Biber Parsed) Data were tagged with the en_docusco_spacy model. R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-docuscope.tabular10M<n<100M0 likes116 downloads2y agoHugging Face13Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes108 downloads3y agoHugging Face14browndw /human-ai-parallel-corpus-spacy Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-spacy.tabular10M<n<100M0 likes105 downloads2y agoHugging Face15HKAllen /cantonese-chinese-parallel-corpus Dataset Summary This dataset consists of parallel sentence pairs in Cantonese and Chinese. It is designed for various tasks, including machine translation. The corpus contains a large number of sentence pairs collected from various domains and most has been improved through manual correction and translation. Languages Cantonese (yue) Simplified Chinese (zh) Dataset Structure Each entry in the dataset is a JSON object containing two fields: "yue" for the… See the full description on the dataset page: https://huggingface.co/datasets/HKAllen/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M3 likes105 downloads2y agoHugging Face16bekan /english_karakalpak_parallel_corpus_v5 English-Karakalpak Parallel Corpus This dataset contains parallel sentences in English and Karakalpak language. It is created to support AI development for the Karakalpak language. Dataset Description English-Karakalpak Parallel Corpus is a high-quality, dynamic dataset containing carefully aligned sentence pairs in English (en) and Karakalpak (kaa). Note: This dataset is updated frequently. New sentence pairs are added on a regular basis to continuously increase… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v5.texttranslation10K<n<100K4 likes104 downloads23d agoHugging Face17ilprl-docse /NepTam-A-Nepali-Tamang-Parallel-Corpus 🧾 NepTam — A Nepali–Tamang Parallel Corpus Dataset Summary NepTam is a high-quality Nepali–Tamang bilingual parallel corpus designed to support research in low-resource neural machine translation (NMT) and linguistic analysis.It contains: 20K gold-standard human-translated sentence pairs, and 80K synthetic pairs generated using the NLLB-200 model fine-tuned on the gold corpus. Each entry includes linguistic metadata such as sentence type, tense, and polarity… See the full description on the dataset page: https://huggingface.co/datasets/ilprl-docse/NepTam-A-Nepali-Tamang-Parallel-Corpus.texttranslation10K<n<100K1 likes102 downloads11mo agoHugging Face18browndw /coca-ai-parallel-corpus-biber COCA-AI Parallel Corpus (Biber Parsed) R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles@misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-biber.tabular10K<n<100K0 likes96 downloads2y agoHugging Face19adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes96 downloads27d agoHugging Face20carles-undergrad-thesis /msmarco-corpus-en-id-parallel-sentences Dataset Card for "msmarco-corpus-en-id-parallel-sentences" More Information needed text1M<n<10M0 likes95 downloads3y agoHugging Face21projecte-aina /CA-ZH_Parallel_Corpus Dataset Card for CA-ZH Parallel Corpus Dataset Summary The CA-ZH Parallel Corpus is a Catalan-Chinese dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Chinese and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The sentences included in the… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-ZH_Parallel_Corpus.texttranslation10M<n<100M2 likes95 downloads1y agoHugging Face22BSC-LT /Catalan-Aranese_Parallel_Corpus Dataset Card for Catalan-Aranese Parallel Corpus Dataset Summary A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.texttranslation100K<n<1M2 likes95 downloads8mo agoHugging Face23sarjukesumo /quran-parallel-corpus Quran Parallel Corpus Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs). Stats Total verses: 6236 Languages: Arabic, English, Indonesian Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian Formats: JSONL, CSV, Parquet Structure Each verse record contains: Field Description surah_number Chapter (1–114) surah_name_arabic Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.tabulartranslation100K<n<1M0 likes91 downloads2mo agoHugging Face24projecte-aina /CA-GL_Parallel_Corpus Dataset Card for CA-GL Parallel Corpus Dataset Description Dataset Summary The CA-GL Parallel Corpus is a Catalan-Galician synthetic dataset parallel sentences created to support the use of co-official languages from Spain, such as Catalan and Galician, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Galician and Catalan in any direction… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-GL_Parallel_Corpus.texttranslation10M<n<100M1 likes84 downloads1y agoHugging Face25Oumar199 /French_Wolof_Various_Parallel_Corpustexttranslation1K<n<10K3 likes82 downloads2y agoHugging Face26projecte-aina /ES-OC_Parallel_Corpus Dataset Card for ES-OC Parallel Corpus Dataset Summary The ES-OC Parallel Corpus is a Spanish-Aranese dataset created to support the use of under-resourced languages from Spain, such as Aranese, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Aranese and Spanish in any direction, as well as Multilingual Machine Translation models. Languages… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-OC_Parallel_Corpus.texttranslation100K<n<1M2 likes71 downloads1y agoHugging Face27PrinceAlhassanNasamu /kusaal-english-parallel-corpus Kusaal-English Parallel Corpus The first open parallel corpus for Kusaal — a Gur language spoken by ~400,000 people in northern Ghana and parts of Burkina Faso. Kusaal has no entry in Google Translate, no presence in Meta's NLLB-200, and no prior open NLP dataset. This corpus was assembled from scratch by a native Kusaal speaker from Bawku, Ghana, and used to train the first open-source Kusaal-English machine translation model: PrinceAlhassanNasamu/kusaal-nllb-600M.… See the full description on the dataset page: https://huggingface.co/datasets/PrinceAlhassanNasamu/kusaal-english-parallel-corpus.texttranslation10K<n<100K2 likes70 downloads3mo agoHugging Face28bekan /english_karakalpak_parallel_corpus_v10 English-Karakalpak Parallel Corpus v10.0 Dataset Description English-Karakalpak Parallel Corpus v10.0 is a high-quality, finalized parallel dataset containing over 50,667 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is structurally optimized to support and accelerate the development of Neural Machine Translation (NMT) systems. Language(s): English (en), Karakalpak (kaa) Format: CSV (Comma-Separated Values) License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v10.texttranslation10K<n<100K0 likes70 downloads23d agoHugging Face29aramasethu /arabic_parallel_corpustext10K<n<100K0 likes69 downloads3y agoHugging Face30kalixlouiis /HFcourse-english-burmese-parallel-corpus HFcourse-English-Burmese-Parallel-Corpus Dataset Description Dataset Summary The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.texttranslation1K<n<10K12 likes69 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.