Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cloverx-id /lumi-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/lumi-repository-parallel-en-id-corpus.tabulartranslation100M<n<1B2 likes2.4k downloads1d agoHugging Face02browndw /human-ai-parallel-corpus Human-AI Parallel English Corpus (HAP-E) 🙃 Purpose The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs). The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words. Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.texttext-classification10K<n<100K5 likes1.7k downloads2y agoHugging Face03mteb /english-danish-parallel-corpus DanishMedicinesAgencyBitextMining An MTEB dataset Massive Text Embedding Benchmark A Bilingual English-Danish parallel corpus from The Danish Medicines Agency. Task category t2t Domains Medical, Written Reference https://sprogteknologi.dk/dataset/bilingual-english-danish-parallel-corpus-from-the-danish-medicines-agency How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/english-danish-parallel-corpus.texttranslation10K<n<100K0 likes229 downloads1y agoHugging Face04browndw /human-ai-parallel-corpus-biber Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.tabular10K<n<100K0 likes184 downloads2y agoHugging Face05adiga-ai /circassian-parallel-corpus Circassian-Russian Parallel Corpus v1.0 This is a high-quality dataset containing over 330,000 parallel text pairs for machine translation between Russian and the Circassian language in its two literary dialects: East Circassian (Kabardian, kbd) and West Circassian (Adyghe, ady). About Circassian Circassian is an indigenous language of the Northwest Caucasus region. The language is notable for its complex phonological system (featuring 50+ consonants)… See the full description on the dataset page: https://huggingface.co/datasets/adiga-ai/circassian-parallel-corpus.texttranslation100K<n<1M5 likes171 downloads1y agoHugging Face06projecte-aina /ES-AN_Parallel_Corpus Dataset Card for ES-AN Parallel Corpus Dataset Summary The ES-AN Parallel Corpus is a Spanish-Aragonese dataset created to support the use of under-resourced languages from Spain, such as Aragonese, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Aragonese and Spanish in any direction, as well as Multilingual Machine Translation models.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-AN_Parallel_Corpus.texttranslation10K<n<100K2 likes132 downloads1y agoHugging Face07browndw /human-ai-parallel-corpus-spacy Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-spacy.tabular10M<n<100M0 likes117 downloads2y agoHugging Face08browndw /human-ai-parallel-corpus-docuscope COCA-AI Parallel Corpus (Biber Parsed) Data were tagged with the en_docusco_spacy model. R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-docuscope.tabular10M<n<100M0 likes92 downloads2y agoHugging Face09browndw /coca-ai-parallel-corpus-biber COCA-AI Parallel Corpus (Biber Parsed) R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles@misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-biber.tabular10K<n<100K0 likes90 downloads2y agoHugging Face10BSC-LT /Catalan-Aranese_Parallel_Corpus Dataset Card for Catalan-Aranese Parallel Corpus Dataset Summary A bilingual parallel corpus for the low-resource language pair Catalan-Aranese. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic Catalan translations generated from… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Catalan-Aranese_Parallel_Corpus.texttranslation100K<n<1M2 likes90 downloads8mo agoHugging Face11projecte-aina /CA-ZH_Parallel_Corpus Dataset Card for CA-ZH Parallel Corpus Dataset Summary The CA-ZH Parallel Corpus is a Catalan-Chinese dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Chinese and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The sentences included in the… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-ZH_Parallel_Corpus.texttranslation10M<n<100M2 likes89 downloads1y agoHugging Face12sarjukesumo /quran-parallel-corpus Quran Parallel Corpus Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs). Stats Total verses: 6236 Languages: Arabic, English, Indonesian Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian Formats: JSONL, CSV, Parquet Structure Each verse record contains: Field Description surah_number Chapter (1–114) surah_name_arabic Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.tabulartranslation100K<n<1M0 likes88 downloads2mo agoHugging Face13projecte-aina /CA-GL_Parallel_Corpus Dataset Card for CA-GL Parallel Corpus Dataset Description Dataset Summary The CA-GL Parallel Corpus is a Catalan-Galician synthetic dataset parallel sentences created to support the use of co-official languages from Spain, such as Catalan and Galician, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Galician and Catalan in any direction… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-GL_Parallel_Corpus.texttranslation10M<n<100M1 likes86 downloads1y agoHugging Face14adeshkin /khakas-russian-parallel-corpus Khakas-Russian Parallel Corpus The creation of this dataset is aimed at supporting the development of natural language processing (NLP) tools and machine translation for the Khakas language, which is classified as a "Definitely Endangered" language. By providing high-quality parallel data, this project helps preserve the linguistic heritage of the Khakas people. Dataset Overlap: The Khakas sentences in this corpus do not overlap with those in the Khakas… See the full description on the dataset page: https://huggingface.co/datasets/adeshkin/khakas-russian-parallel-corpus.texttranslation100K<n<1M2 likes76 downloads1mo agoHugging Face15projecte-aina /CA-DE_Parallel_Corpus Dataset Card for CA-DE Parallel Corpus Dataset Summary The CA-DE Parallel Corpus is a Catalan-German dataset of parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between German and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The sentences included in the… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-DE_Parallel_Corpus.texttranslation1M<n<10M0 likes75 downloads1y agoHugging Face16aramasethu /arabic_parallel_corpustext10K<n<100K0 likes69 downloads3y agoHugging Face17aykgeh /Ekegusii-English-Kiswahili-Parallel-Corpus Ekegusii - English - Kiswahili Parallel Master Corpus This repository contains the official consolidated, cleaned, and deduplicated Multilingual Parallel Corpus for Ekegusii (guz), Kiswahili (sw), and English (en) machine translation research, with a focus on Public Service Announcements (PSAs) in Kenya. Dataset Summary Total Records: 149,849 unique parallel alignment concepts. Languages: English, Kiswahili (Swahili), Ekegusii (Gusii). Public Service… See the full description on the dataset page: https://huggingface.co/datasets/aykgeh/Ekegusii-English-Kiswahili-Parallel-Corpus.texttranslation100K<n<1M1 likes66 downloads1mo agoHugging Face18carles-undergrad-thesis /msmarco-corpus-en-id-parallel-sentences Dataset Card for "msmarco-corpus-en-id-parallel-sentences" More Information needed text1M<n<10M0 likes64 downloads3y agoHugging Face19projecte-aina /ES-OC_Parallel_Corpus Dataset Card for ES-OC Parallel Corpus Dataset Summary The ES-OC Parallel Corpus is a Spanish-Aranese dataset created to support the use of under-resourced languages from Spain, such as Aranese, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Aranese and Spanish in any direction, as well as Multilingual Machine Translation models. Languages… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-OC_Parallel_Corpus.texttranslation100K<n<1M2 likes61 downloads1y agoHugging Face20projecte-aina /CA-IT_Parallel_Corpus Dataset Card for CA-IT Parallel Corpus Dataset Summary The CA-IT Parallel Corpus is a Catalan-Italian dataset parallel sentences created to support Catalan in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Italian and Catalan in any direction, as well as Multilingual Machine Translation models. Languages The sentences included in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-IT_Parallel_Corpus.texttranslation1M<n<10M0 likes59 downloads1y agoHugging Face21raptorkwok /cantonese-chinese-parallel-corpus Cantonese-Written Chinese Parallel Corpus (CCPC) About the Dataset Data Splits Training Data (train): 160,000 Sentence Pairs Validation Data (validation): 20,000 Sentence Pairs Test Data (test): 5,461 Sentence Pairs Languages Cantonese (yue) Traditional Chinese (zh-TW) Original Data Structure JSON lines consisting of yue, zh and ref fields. Data Source Apart from the data from our first generation of CCPC, there are… See the full description on the dataset page: https://huggingface.co/datasets/raptorkwok/cantonese-chinese-parallel-corpus.texttranslation100K<n<1M4 likes54 downloads5mo agoHugging Face22BSC-LT /Spanish-Valencian_Catalan_Parallel_Corpus Dataset Card for Spanish-Valencian Catalan Parallel Corpus Dataset Summary A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.texttranslation1M<n<10M2 likes54 downloads7mo agoHugging Face23projecte-aina /ES-AST_Parallel_Corpus Dataset Card for ES-AST Parallel Corpus Dataset Summary The ES-AST Parallel Corpus is a Spanish-Asturian dataset created to support the use of under-resourced languages from Spain, such as Asturian, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Asturian and Spanish in any direction, as well as Multilingual Machine Translation models.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/ES-AST_Parallel_Corpus.texttranslation100K<n<1M3 likes53 downloads1y agoHugging Face24IbrahimAmin /arz-en-parallel-corpus Egyptian Arabic - English Parallel Corpus 🇪🇬✨🇬🇧 Dataset Description This dataset is a cleaned and filtered merge of multiple Egyptian Arabic - English parallel corpora, containing ~27,000 aligned sentence pairs. It’s designed for researchers and developers working on machine translation, speech translation, and other NLP tasks involving Egyptian Arabic and English. Sources 📚 This dataset integrates and refines data from the following publicly available… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/arz-en-parallel-corpus.texttranslation10K<n<100K3 likes50 downloads1y agoHugging Face25browndw /human-ai-parallel-corpus-2-emotionstabular1M<n<10M0 likes50 downloads8mo agoHugging Face26kingkaung /islamqainfo_parallel_corpus Dataset Card for IslamQA Info Parallel Corpus Dataset Description The IslamQA Info Parallel Corpus is a multilingual dataset derived from the IslamQA repository. It contains curated question-and-answer pairs across 17 languages, making it a valuable resource for multilingual and cross-lingual natural language processing (NLP) tasks. The dataset has been created over nearly three decades (since 1997) by Sheikhul Islam Muhammad Saalih al-Munajjid and his team. Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/islamqainfo_parallel_corpus.texttable-question-answering10K<n<100K2 likes48 downloads2y agoHugging Face27tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes45 downloads1y agoHugging Face28projecte-aina /CA-EU_Parallel_Corpus Dataset Card for CA-EU Parallel Corpus Dataset Summary The CA-EU Parallel Corpus is a Catalan-Basque synthetic dataset of parallel sentences created to support the use of co-official languages from Spain, such as Catalan and Basque, in NLP tasks, specifically Machine Translation. Supported Tasks and Leaderboards The dataset can be used to train Bilingual Machine Translation models between Basque and Catalan in any direction, as well as Multilingual Machine… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CA-EU_Parallel_Corpus.texttranslation10M<n<100M0 likes44 downloads2y agoHugging Face29browndw /coca-ai-parallel-corpus-spacy Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-spacy.tabular10M<n<100M0 likes43 downloads2y agoHugging Face30guiti999 /cantonese-traditional-chinese-parallel-corpus-gen3 Cantonese-Written Chinese Parallel Dataset (3rd Generation) About the Dataset Data Splits Training Data (train): 160,000 Sentence Pairs Validation Data (validation): 20,000 Sentence Pairs Test Data (test): 5,461 Sentence Pairs Languages Cantonese (yue) Traditional Chinese (zh-TW) Original Data Structure JSON lines consisting of yue, zh and ref fields. Data Source LIHKG HKCancor Cantonse-Mandarin Translations and various… See the full description on the dataset page: https://huggingface.co/datasets/guiti999/cantonese-traditional-chinese-parallel-corpus-gen3.texttranslation100K<n<1M0 likes43 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.