Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cloverx-id /lumi-repository-parallel-en-id-corpusWe are currently developing new version of LMSE translation scoring model and processing additional data sources. We estimate the dataset will expand, with significantly improved quality.(Delayed..) A score of 55% and above indicates high-quality translation pairs, even if the first version of the model we developed gave them such a score. We will try to release a newer model in the future with better quality and consistently fast scoring speeds, and release it to the public once we decide… See the full description on the dataset page: https://huggingface.co/datasets/cloverx-id/lumi-repository-parallel-en-id-corpus.tabulartranslation100M<n<1B2 likes2.4k downloads1d agoHugging Face02browndw /human-ai-parallel-corpus-biber Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-biber.tabular10K<n<100K0 likes184 downloads2y agoHugging Face03browndw /human-ai-parallel-corpus-spacy Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-spacy.tabular10M<n<100M0 likes117 downloads2y agoHugging Face04browndw /human-ai-parallel-corpus-docuscope COCA-AI Parallel Corpus (Biber Parsed) Data were tagged with the en_docusco_spacy model. R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do… See the full description on the dataset page: https://huggingface.co/datasets/browndw/human-ai-parallel-corpus-docuscope.tabular10M<n<100M0 likes92 downloads2y agoHugging Face05browndw /coca-ai-parallel-corpus-biber COCA-AI Parallel Corpus (Biber Parsed) R users can import the data directly using r-polars: library(polars) df <- pl$read_parquet('hf://datasets/browndw/coca-ai-parallel-corpus-biber/**/*.parquet') df <- df$to_data_frame() Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles@misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-biber.tabular10K<n<100K0 likes90 downloads2y agoHugging Face06sarjukesumo /quran-parallel-corpus Quran Parallel Corpus Verse-aligned Quran parallel corpus — Arabic (Uthmani), English (Sahih International), and Indonesian (Ministry of Religious Affairs). Stats Total verses: 6236 Languages: Arabic, English, Indonesian Translation pairs: Arabic↔English, Arabic↔Indonesian, English↔Indonesian Formats: JSONL, CSV, Parquet Structure Each verse record contains: Field Description surah_number Chapter (1–114) surah_name_arabic Arabic surah… See the full description on the dataset page: https://huggingface.co/datasets/sarjukesumo/quran-parallel-corpus.tabulartranslation100K<n<1M0 likes88 downloads2mo agoHugging Face07browndw /human-ai-parallel-corpus-2-emotionstabular1M<n<10M0 likes50 downloads8mo agoHugging Face08tunis-ai /tunisian-msa-parallel-corpus Dataset Description This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models. The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.tabulartranslation1K<n<10K0 likes45 downloads1y agoHugging Face09browndw /coca-ai-parallel-corpus-spacy Citation If you use the corpus as part of your research, please cite: Do LLMs write like humans? Variation in grammatical and rhetorical styles @misc{reinhart2024llmswritelikehumans, title={Do LLMs write like humans? Variation in grammatical and rhetorical styles}, author={Alex Reinhart and David West Brown and Ben Markey and Michael Laudenbach and Kachatad Pantusen and Ronald Yurko and Gordon Weinberg}, year={2024}, eprint={2410.16107}… See the full description on the dataset page: https://huggingface.co/datasets/browndw/coca-ai-parallel-corpus-spacy.tabular10M<n<100M0 likes43 downloads2y agoHugging Face10nickoo004 /kaa-parallel-corpus Kaa Karakalpak-English Parallel Corpus (FineTranslations) 📌 Overview This repository contains a high-quality, curated parallel corpus for the Karakalpak (kaa) language, paired with English (en). Karakalpak is a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan. This dataset is a specialized subset extracted from the massive HuggingFaceFW/finetranslations project. The goal of this repo is to provide a dedicated and easy-to-access resource… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/kaa-parallel-corpus.tabulartranslation10K<n<100K0 likes42 downloads6mo agoHugging Face11gvij /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus-formattedtabular1M<n<10M0 likes39 downloads2y agoHugging Face12LocalDoc /en-az-opus-filtered-parallel-corpus Filtered EN-AZ OPUS Parallel Corpus English–Azerbaijani parallel sentences pooled from OPUS corpora and filtered with a two-stage quality-estimation pipeline. Pipeline LaBSE cross-lingual cosine similarity (kept the well-aligned pairs). COMET-Kiwi (Unbabel/wmt22-cometkiwi-da) reference-free QE on the survivors. Exact-pair deduplication. Effective minimums in this release: LaBSE ≥ 0.900, COMET-Kiwi ≥ 0.900. Columns en_text — English (source)… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/en-az-opus-filtered-parallel-corpus.tabulartranslation100K<n<1M0 likes23 downloads4mo agoHugging Face13tunis-ai /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K2 likes22 downloads1y agoHugging Face14sarjukesumo /hadith-parallel-corpustabular1M<n<10M0 likes22 downloads2mo agoHugging Face15ArabicNLPWorld /arabic-russian-parallel-corpusgated Arabic‑Russian Parallel Corpus A parallel corpus for Arabic–Russian language pairs. Each record contains an Arabic sentence/phrase, its Russian translation, and the source of the pair.The dataset has been cleaned, deduplicated, and source names normalized to lowercase. 📊 Dataset Statistics Overview Metric Value Total entries 116,393 Unique Arabic strings 116,124 Unique Russian strings 116,152 Unique sources 6 Data… See the full description on the dataset page: https://huggingface.co/datasets/ArabicNLPWorld/arabic-russian-parallel-corpus.tabular100K<n<1M0 likes18 downloads4mo agoHugging Face16hbenayed /tunisian-msa-parallel-corpus-evaluated Dataset Description This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation. Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.tabulartranslation1K<n<10K0 likes12 downloads6mo agoHugging Face17laurentiubp /CA-EN_Parallel_Corpustabular10K<n<100K0 likes6 downloads2y agoHugging Face18stukenov /ekitil-corpus-parallel-kkru-v1gatedtabular100K<n<1M0 likes6 downloads7mo agoHugging Face19burkimbia /moore-parallel-corpusgated Mooré–French parallel corpus French–Mooré sentence and term pairs for machine translation, built by bia-datasets-text from 18 cooked sources: cleaned, orthography-normalized, annotated, filtered and deduplicated, with the frozen BurkimbIA MT benchmark's source texts excluded from every split. from datasets import load_dataset ds = load_dataset("burkimbia/moore-parallel-corpus", revision="v0.9.0") # this release; omit revision for the latest Usage and license… See the full description on the dataset page: https://huggingface.co/datasets/burkimbia/moore-parallel-corpus.tabulartranslation100K<n<1M0 likes7h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.