datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UN_Historical_PDF_Article_Text_Corpus
python
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train")
or
dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest")
lang_list = ["ar", "en", "es", "fr", "ru", "zh"]
for row in dataset:
# 获取pdf文章内容
for lang in lang_list:
# type == str
lang_match_file_content = row[lang]
# 如果按页分割
lang_match_file_pages_content = lang_match_file_content.split("\n----\n")
ILUR-news-text-classification-corpus
News Texts Dataset
We release a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens).
For more details, refer to the paper.
hindi-english-raw-text-corpus-uncleanedraw-text-corpus
📝 Zomi Raw Text Corpus (Community-Contributed)
The Zomi Raw Text Corpus is an open, community-driven collection of unprocessed Zomi-language text.It serves as the foundational dataset for building the full Zomi NLP ecosystem, including tokenizers, language models, ASR/TTS systems, and downstream tasks.
This dataset is intentionally raw — no normalization, deduplication, or cleaning is applied.Cleaned and task-specific datasets will be released separately.
📦 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/raw-text-corpus.persian-text-corpus
Persian Corpus (Merged)
Dataset Summary
Persian Corpus (Merged) is a large-scale, Persian corpus meticulously aggregated from multiple high-quality Persian datasets available on the Hugging Face Hub. Designed to advance Persian NLP research and applications, this corpus consolidates diverse textual sources into a single resource, providing researchers and developers with a robust foundation for training and evaluating language models.
Why Use This… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-text-corpus.ru-text-corpus
Description
798k deduplicated Russian documents (1.6B tokens) from FineWeb-2 (rus_Cyrl), ru-StackOverflow, Pikabu, Habr, Russian Wikipedia and news. Markdown- and code-bearing (StackOverflow/Habr/Pikabu are text_markdown). Filtered to >=200 chars and >=30% Cyrillic. Columns: text, src.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/PotatoHD/ru-text-corpus.Nepali-Text-Corpus
Nepali Text Corpus
Overview
Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a
diverse range of text types, including news articles, blogs, and more, making it an invaluable
resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP)
and computational linguistics.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/IRIIS-RESEARCH/Nepali-Text-Corpus.nepali-text-corpus-64
Nepali Text Dataset
Overview
The Nepali Text Dataset is a comprehensive collection of approximately 6.4 million articles in the
Nepali language. This dataset encompasses a diverse range of text types, including news articles, blogs,
and more, making it an invaluable resource for researchers, developers, and enthusiasts
in the fields of Natural Language Processing (NLP) and computational linguistics.
Dataset Details
Total Articles: ~6.4 million
Language:… See the full description on the dataset page: https://huggingface.co/datasets/mridul3301/nepali-text-corpus-64.roblox_docs_corpus_text
[!Note]
Last collected: 2026-10-01 15:11
Contains 1789 contents
parla_text_corpus
ParlaTextCorpus
Spoken text corpus for Catalan. Derived and cleaned from three sources. OpenSubtitles, Tv3Parla and Festcat.
parla_text_corpusblk-text-corpus
Verified Pa'O (Blk) Text Corpus - Pa'O Digital Hub
Dataset Summary
This is the official, verified parallel dataset for the Pa'O language (ISO 639-3: blk) and Burmese (Myanmar) translations, published by Pa'O Digital Hub.
The corpus is systematically collected, reviewed, standardized, and verified through the established linguistic and editorial workflow of Pa'O Digital Hub. The Pa'O sentences are based on authentic language usage by Pa'O native speakers and are… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/blk-text-corpus.khmer-text-corpus
Khmer + English Mixed Text Corpus
A cleaned, deduplicated Khmer text corpus with naturally-occurring Khmer/English
code-switching (English tech & finance terms, Latin script, and digits embedded in Khmer
text). It is the training data for the
Panhapich/khmer-sp-8k SentencePiece tokenizer and the
text-only warm-start of a shared Khmer diffusion decoder.
Dataset summary
4,893,739 sentences, one per line, UTF-8, deduplicated and shuffled
(fixed seed 42… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-text-corpus.burmese-text-corpus
Burmese Text Corpus For Natural Language Processing
🎫 Choose your language: 🌏 English Version | 🇲🇲 မြန်မာဗားရှင်း
🌏 English Version
This dataset is a specifically curated text corpus for the Burmese language. It is intended to support Natural Language Processing (NLP) tasks, language model training, and research related to the Burmese language.
1. About the Dataset
The primary goal of creating this burmese-text-corpus dataset is to address the scarcity of… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/burmese-text-corpus.matina_persian_text_corpus
Matina: A Large-Scale 73B Token Persian Text Corpus
Text corpora are essential for training models used in tasks like summarization, translation, and large language models (LLMs).
While various efforts have been made to collect monolingual and multilingual datasets in many languages,
Persian has often been underrepresented due to limited resources for data collection and preprocessing.
Existing Persian datasets are typically small and lack content diversity, consisting mainly of… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/matina_persian_text_corpus.Cantonese-Japanese-Machine-Translation-Corpus-text
PangeanicYueJa - Cantonese Japanese Parallel Corpus
PangeanicYueJa is a Cantonese-Japanese parallel corpus designed for machine translation, multilingual large language model (LLM) training, cross-lingual NLP research, retrieval-augmented generation (RAG), bilingual embeddings, instruction tuning, and multilingual AI systems.
This release contains 55,000 Cantonese-Japanese sentence pairs sampled from a larger corpus of approximately 3.08 million parallel sentence pairs. For the… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Cantonese-Japanese-Machine-Translation-Corpus-text.video-text-corpus
Code Video Text Data Notes
Dataset summary
A documented Code data-preparation workflow for Video Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/dewozniak/video-text-corpus.video-text-corpus
Movie Posters Video Text Data Notes
Dataset summary
A documented Movie Posters data-preparation workflow for Video Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/thomaswardzij/video-text-corpus.video-text-corpus
Math Video Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Math work with Video Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Oscarbaek/video-text-corpus.video-text-corpus51
Comics Video Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Comics work with Video Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/subramanianno/video-text-corpus51.Tumbuka_Text_Corpus_Translated_Gutenberg
Tumbuka Text Corpus - Translated Gutenberg
Dataset Description
This dataset contains a large-scale collection of Tumbuka text, primarily consisting of machine-translated literary works from the Project Gutenberg library. It is designed to support Natural Language Processing (NLP) research for Tumbuka, a Bantu language spoken in Malawi, Zambia, and Tanzania.
Dataset Summary
Language: Tumbuka (tum)
Source: Project Gutenberg
Content: Translated… See the full description on the dataset page: https://huggingface.co/datasets/Mwanzau/Tumbuka_Text_Corpus_Translated_Gutenberg.Nepali-Extended-Text-Corpus
Dataset Card for "Nepali-Extended-Text-Corpus"
More Information needed
text-tabular-corpus
Gaming Text Tabular Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Text Tabular data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/heitorlr/text-tabular-corpus.vi-text_corpus-dantri.com.vn-splittedvideo-text-corpus
Geology Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Geology tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card… See the full description on the dataset page: https://huggingface.co/datasets/FELIXFF92/video-text-corpus.image-text-corpus
Speech Image Text Data Notes
Dataset summary
A documented Speech data-preparation workflow for Image Text records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/jolawal1995/image-text-corpus.CiviVox-English-Swahili-text-translation-corpusadaption-amharic-text-corpus
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-amharic_text_corpus
This dataset comprises over 700,000 Amharic text documents formatted as line-delimited JSON, covering diverse topics such as history, religion, politics, and product descriptions. Each entry contains a single string field with native Amharic content, including some samples with mixed languages or placeholder values. It is designed for text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-amharic-text-corpus.product_review_text_corpus
Academic Product Review Text Corpus
Academic corpus of product review texts assembled by a university research group.
Terms
Released under the Creative Commons Attribution 4.0 International license (CC-BY-4.0).
text-tabular-corpus
Agriculture Text Tabular Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Agriculture work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/cody-jones/text-tabular-corpus.
