datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lumi-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..)
A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/lumi-repository-parallel-en-id-corpus.human-ai-parallel-corpus-biber
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.human-ai-parallel-corpus-spacy
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-spacy.human-ai-parallel-corpus-docuscope
COCA-AI Parallel Corpus (Biber Parsed)
Data were tagged with the en_docusco_spacy model.
R users can import the data directly using r-polars:
library(polars)
df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet')
df <- df$to_data_frame()
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-docuscope.coca-ai-parallel-corpus-biber
COCA-AI Parallel Corpus (Biber Parsed)
R users can import the data directly using r-polars:
library(polars)
df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet')
df <- df$to_data_frame()
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-biber.quran-parallel-corpus
Quran Parallel Corpus
Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs).
Stats
Total verses: 6236
Languages: Arabic, English, Indonesian
Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian
Formats: JSONL, CSV, Parquet
Structure
Each verse record contains:
Field
Description
surah_number
Chapter (1–114)
surah_name_arabic
Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.human-ai-parallel-corpus-2-emotionstunisian-msa-parallel-corpus
Dataset Description
This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models.
The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.coca-ai-parallel-corpus-spacy
Citation
If you use the corpus as part of your research, please cite:
Do LLMs write like humans? Variation in grammatical and rhetorical styles
@misc{reinhart2024llmswritelikehumans,
title={Do LLMs write like humans? Variation in grammatical and rhetorical styles},
author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg},
year={2024},
eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-spacy.kaa-parallel-corpus
Kaa Karakalpak-English Parallel Corpus (FineTranslations)
📌 Overview
This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan.
This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.JParaCrawl-Filtered-English-Japanese-Parallel-Corpus-formatteden-az-opus-filtered-parallel-corpus
Filtered EN-AZ OPUS Parallel Corpus
English–Azerbaijani parallel sentences pooled from OPUS corpora and filtered with a
two-stage quality-estimation pipeline.
Pipeline
LaBSE cross-lingual cosine similarity (kept the well-aligned pairs).
COMET-Kiwi (Unbabel/wmt22-cometkiwi-da) reference-free QE on the survivors.
Exact-pair deduplication.
Effective minimums in this release: LaBSE ≥ 0.900, COMET-Kiwi ≥ 0.900.
Columns
en_text — English (source)… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/en-az-opus-filtered-parallel-corpus.tunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.hadith-parallel-corpusarabic-russian-parallel-corpus
Arabic‑Russian Parallel Corpus
A parallel corpus for Arabic–Russian language pairs. Each record contains an Arabic sentence/phrase, its Russian translation, and the source of the pair.The dataset has been cleaned, deduplicated, and source names normalized to lowercase.
📊 Dataset Statistics
Overview
Metric
Value
Total entries
116,393
Unique Arabic strings
116,124
Unique Russian strings
116,152
Unique sources
6
Data… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-parallel-corpus.tunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.CA-EN_Parallel_Corpusekitil-corpus-parallel-kkru-v1moore-parallel-corpus
Mooré–French parallel corpus
French–Mooré sentence and term pairs for machine translation, built by bia-datasets-text from 18 cooked sources: cleaned, orthography-normalized, annotated, filtered and deduplicated, with the frozen BurkimbIA MT benchmark's source texts excluded from every split.
from datasets import load_dataset
ds = load_dataset("burkimbia/moore-parallel-corpus", revision="v0.9.0") # this release; omit revision for the latest
Usage and license… See the full description on the dataset page: https://huggingface.co/datasets/burkimbia/moore-parallel-corpus.
