parallel-corpus
lumi-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..)
A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/lumi-repository-parallel-en-id-corpus.human-ai-parallel-corpus
Human-AI Parallel English Corpus (HAP-E) 🙃
Purpose
The HAP-E corpus is designed for comparisions of the writing produced by humans and the writing produced by large language models (LLMs).
The corpus was created by seeding an LLM with an approximately 500-word chunk of human-authored text and then prompting the model to produce an additional 500 words.
Thus, a second 500-word chunk of human-authored text (what actually comes next in the original text) can be compared to… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus.Egyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.UFAL_Parallel_Corpus_of_North_Levantine_1.0
[!NOTE]
Dataset origin: https://zenodo.org/records/4012218
UFAL Parallel Corpus of North Levantine 1.0
March 10, 2023
Authors
Shadi Saleh <saleh@ufal.mff.cuni.cz>
Hashem Sellat <sellat@ufal.mff.cuni.cz>
Mateusz Krubiński <krubinski@ufal.mff.cuni.cz>
Adam Posppíšil <adam.pospisil@ff.cuni.cz>
Petr Zemánek <petr.zemanek@ff.cuni.cz>
Pavel Pecina <pecina@ufal.mff.cuni.cz>
Overview
This is the first release of the UFAL Parallel Corpus of North Levantine… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/UFAL_Parallel_Corpus_of_North_Levantine_1.0.english-danish-parallel-corpus
DanishMedicinesAgencyBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
A Bilingual English-Danish parallel corpus from The Danish Medicines Agency.
Task category
t2t
Domains
Medical, Written
Reference
https://sprogteknologi.dk/dataset/bilingual-english-danish-parallel-corpus-from-the-danish-medicines-agency
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/english-danish-parallel-corpus.african-language-parallel-corpus
African Language Parallel Corpus
Human-created, human-validated parallel sentence pairs for three African languages,
released openly by Okwu. Version 1.5.
Dataset summary
A parallel corpus of everyday-register sentence pairs for Yorùbá, Swahili, and
Nigerian Pidgin, each paired with English. The core is derived from NKENNE's own
language-learning curriculum — content authored and reviewed by native-speaker educators —
supplemented for Swahili with public-domain… See the full description on the dataset page: https://huggingface.co/datasets/Okwu/african-language-parallel-corpus.
