datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release:
HPLT3.0
We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0.
This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to our website and our pre-print.
The Cleaned variant of HPLT Datasets v2.0
This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.DocHPLT
DocHPLT: A Massively Multilingual Document-Level Translation Dataset
Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.hplt2_edu_scores
HPLT2-Edu-scores
Dataset summary
HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance.
For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.hpltv2-llama33-edu-annotation
HPLT version 2.0 educational annotations
This dataset contains annotations derived from HPLT v2 cleaned samples.
There are 500,000 annotations for each language if the source contains at least 500,000 samples.
We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier.
Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.hplt2.0_cleaned_urls
Dataset Card for hplt2.0_cleaned_urls
This dataset provides the URLs and top-level domains associated with training records in HPLT/HPLT2.0_cleaned. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt2.0_cleaned_urls.hplt3_edu_scores
HPLT3-Edu-scores
Dataset summary
HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages.
Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction.
HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings.
For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.indic-hplt-v2
Indic HPLT v2
A multilingual pretraining corpus of 34,605,630 documents (~25.5B estimated tokens, ~218 GB raw JSONL) across 13 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data.
This is the larger successor to Indic HPLT v1 (9.8M docs, 11 languages). Compared to v1, this release adds 3 new Indic languages (Nepali, Odia, Assamese) and ~3.5× more documents overall.
Quick Start
from datasets import load_dataset
# Full training… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v2.hplt3_domainshplt2_embeddings
HPLT2-embeddings
Dataset summary
HPLT2-embeddings is an extension of the HPLT2 dataset, annotated with document-level Snowflake's Arctic-embed-m-v2.0 embeddings for 35 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research.
Snowflake-arctic-embed-m-v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document.
The… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_embeddings.indic-hplt-v1
Indic HPLT v1
A multilingual pretraining corpus of 9,836,075 documents (~8.4B estimated tokens) across 10 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data.
Quick Start
from datasets import load_dataset
# Full training split
ds = load_dataset("ashtok897/indic-hplt-v1", split="train")
# Filter by language
hi_ds = ds.filter(lambda x: x["lang"] == "hi")
# Streaming (recommended for large-scale use)
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v1.african-languages-hplt-filtered
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.european-hplt-v1
European HPLT v1
A multilingual pretraining corpus of 45,031,396 documents (~50.9B estimated tokens, ~190 GB raw JSONL) across 41 European languages, built from HPLT Monolingual v3 high-quality web crawl data.
The corpus spans Germanic, Romance, Slavic, Celtic, Baltic, Finno-Ugric, Greek, and other European language families. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution).
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/european-hplt-v1.hplt_bnhplt
HPLT Indic Language Corpus
This repository contains selected Indic-language data from the HPLT Monolingual Dataset 3.0, prepared and hosted by CoRover for large-scale generative language model pretraining and multilingual NLP research.
The dataset contains raw HPLT data for multiple Indian languages and scripts.
Languages
The repository currently contains the following language/script datasets:
Language
Language Code
Script
Directory
Bengali
ben
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/CoRover/hplt.turkish-hplt2-filteredHPLT3-198-500k
Dataset Card for HPLT3 Multilingual JSONL (Subset)
This card documents the language coverage and document counts for a multilingual dataset built from a subset of HPLT3-style sources.The data are organized as one JSONL file per language–script code (e.g., deu_Latn.jsonl). Each line is one document.
Total documents (lines across all listed files): 51,366,154
Dataset Summary
Format: JSON Lines (.jsonl) — one document per line.
Organization: one file per… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/HPLT3-198-500k.HPLT2.0_cleanedThis is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project.
The source of the data is mostly Internet Archive with some additions from Common Crawl.
For a detailed description of the dataset, please refer to https://hplt-project.org/datasets/v2.0
The Cleaned variant of HPLT Datasets v2.0
This is the cleaned variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here.
The original JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/HPLT2.0_cleaned.HPLT-zhOpenLID-v3
Dataset Description
OpenLID-v3 is an updated version of the OpenLID-v2 dataset (see the CHANGELOG.md).
Repository: https://github.com/hplt-project/openlid
Paper: OpenLID-v3: Improving the Precision of Closely Related Language Identification – An Experience Report
Usage
from datasets import load_dataset
ds = load_dataset('HPLT/OpenLID-v3', split='train')
Dataset Summary
The OpenLID-v3 dataset covers 194 language varieties + not-a-language class… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/OpenLID-v3.HPLT3.0
This is a large-scale collection of web-crawled documents in 198 world languages, produced by the HPLT project.
The source of the data is Internet Archive and Common Crawl.
For a detailed description of this and previous releases by HPLT, please refer to our website.
NB: the HPLT datasets are not hosted on HuggingFace! See download instructions below.
HPLT release v3.0
In July 2025, the European HPLT initiative has completed a new release of its monolingual datasets, offering… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT3.0.HPLT3_DE_0.9_Quantile_Adult_FilteredDocHPLT-v3-clean
DocHPLT v3 cleaned
Cleaned and re-formatted version of DocHPLTv3 ready for LLM training or document-level NMT training.
HPLT_Finnish_fineweb_edu_predictedhplt-filteredhplt-v3-pl-cleaned
HPLT v3 Polish — Cleaned & PII-Gated
Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate.
Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.lumiopen-hpltv2-llama33-edu-annotation-etHPLT_1.2_fi_cleaned2508-datasets-evals
HPLT 3.0: Details on Corpus Comparison Results
Dataset Description
This dataset contains fine-grained results from our HPLT 3.0 release evaluations comparing the new HPLT 3.0 corpora with the previous HPLT 2.0 version, FineWeb2, and MADLAD-400. We pretrain 2.2B Llama-style decoder models on 100B tokens for each selected language and evaluate them using HPLT-E, a multilingual evaluation framework for comprehensive multi-prompt k-shot evaluation across 124 tasks and 500+… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/2508-datasets-evals.hplt-v1.2_urls
Dataset Card for hplt-v1.2_urls
This dataset provides the URLs and top-level domains associated with training records in HPLT v1.2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt-v1.2_urls.hplt3_pol_8_9_10_apt_tokenized_split
