Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B46 likes187k downloads4mo agoHugging Face02HPLT /DocHPLT DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/DocHPLT.texttranslation100M<n<1B20 likes3.9k downloads9mo agoHugging Face03JQL-AI /hplt2_edu_scores HPLT2-Edu-scores Dataset summary HPLT2-JQL-Education is a model-annotated language subset of HPLT2, spanning 35 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_edu_scores.tabulartext-ranking1B<n<10B1 likes3.7k downloads1y agoHugging Face04LumiOpen /hpltv2-llama33-edu-annotation HPLT version 2.0 educational annotations This dataset contains annotations derived from HPLT v2 cleaned samples. There are 500,000 annotations for each language if the source contains at least 500,000 samples. We prompt Llama-3.3-70B-Instruct to score web pages based on their educational value following FineWeb-Edu classifier. Note 1: The dataset contains the prompt (using the first 1500 characters of the text sample), the scores, and the full Llama 3 generation. The column "idx"… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/hpltv2-llama33-edu-annotation.tabular10M<n<100M3 likes1.8k downloads1y agoHugging Face05nhagar /hplt2.0_cleaned_urls Dataset Card for hplt2.0_cleaned_urls This dataset provides the URLs and top-level domains associated with training records in HPLT/HPLT2.0_cleaned. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt2.0_cleaned_urls.text10B<n<100B0 likes1.1k downloads1y agoHugging Face06Eurolingua /hplt3_edu_scores HPLT3-Edu-scores Dataset summary HPLT3-JQL-Education is a model-annotated language subset of HPLT3, spanning 36 languages. Our model-annotations allow for a filtering that achieves higher-quality training outcomes without excessively aggressive data reduction. HPLT3-Edu-scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic-embed-m-v2.0 embeddings. For all training ablations, we used… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/hplt3_edu_scores.tabulartext-ranking1B<n<10B0 likes1k downloads7mo agoHugging Face07ashtok897 /indic-hplt-v2 Indic HPLT v2 A multilingual pretraining corpus of 34,605,630 documents (~25.5B estimated tokens, ~218 GB raw JSONL) across 13 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data. This is the larger successor to Indic HPLT v1 (9.8M docs, 11 languages). Compared to v1, this release adds 3 new Indic languages (Nepali, Odia, Assamese) and ~3.5× more documents overall. Quick Start from datasets import load_dataset # Full training… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v2.tabulartext-generation10M<n<100M3 likes946 downloads4mo agoHugging Face08Eurolingua /hplt3_domainstext1B<n<10B0 likes928 downloads10mo agoHugging Face09JQL-AI /hplt2_embeddings HPLT2-embeddings Dataset summary HPLT2-embeddings is an extension of the HPLT2 dataset, annotated with document-level Snowflake's Arctic-embed-m-v2.0 embeddings for 35 languages, making the dataset useful for a variety of tasks, including document clustering, filtering, and other multilingual research. Snowflake-arctic-embed-m-v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document. The… See the full description on the dataset page: https://huggingface.co/datasets/JQL-AI/hplt2_embeddings.feature-extractionn>1T0 likes801 downloads1y agoHugging Face10ashtok897 /indic-hplt-v1 Indic HPLT v1 A multilingual pretraining corpus of 9,836,075 documents (~8.4B estimated tokens) across 10 Indic languages and English, built from HPLT Monolingual v3 high-quality web crawl data. Quick Start from datasets import load_dataset # Full training split ds = load_dataset("ashtok897/indic-hplt-v1", split="train") # Filter by language hi_ds = ds.filter(lambda x: x["lang"] == "hi") # Streaming (recommended for large-scale use) ds =… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/indic-hplt-v1.tabulartext-generation1M<n<10M4 likes693 downloads4mo agoHugging Face11rufatronics /african-languages-hplt-filtered VelkroLM African Languages Corpus This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative. This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.text1M<n<10M0 likes484 downloads2mo agoHugging Face12ashtok897 /european-hplt-v1 European HPLT v1 A multilingual pretraining corpus of 45,031,396 documents (~50.9B estimated tokens, ~190 GB raw JSONL) across 41 European languages, built from HPLT Monolingual v3 high-quality web crawl data. The corpus spans Germanic, Romance, Slavic, Celtic, Baltic, Finno-Ugric, Greek, and other European language families. Every document has an HPLT WDS quality score of 10 or higher (the top of the quality distribution). Quick Start from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ashtok897/european-hplt-v1.tabulartext-generation10M<n<100M3 likes438 downloads4mo agoHugging Face13aniketsen /hplt_bntext1M<n<10M0 likes429 downloads2y agoHugging Face14CoRover /hplt HPLT Indic Language Corpus This repository contains selected Indic-language data from the HPLT Monolingual Dataset 3.0, prepared and hosted by CoRover for large-scale generative language model pretraining and multilingual NLP research. The dataset contains raw HPLT data for multiple Indian languages and scripts. Languages The repository currently contains the following language/script datasets: Language Language Code Script Directory Bengali ben Bengali… See the full description on the dataset page: https://huggingface.co/datasets/CoRover/hplt.text-generation0 likes370 downloads1mo agoHugging Face15elifozgeyilmaz /turkish-hplt2-filteredtabular10M<n<100M0 likes362 downloads5mo agoHugging Face16Eurolingua /HPLT3-198-500k Dataset Card for HPLT3 Multilingual JSONL (Subset) This card documents the language coverage and document counts for a multilingual dataset built from a subset of HPLT3-style sources.The data are organized as one JSONL file per language–script code (e.g., deu_Latn.jsonl). Each line is one document. Total documents (lines across all listed files): 51,366,154 Dataset Summary Format: JSON Lines (.jsonl) — one document per line. Organization: one file per… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/HPLT3-198-500k.text-generation10M<n<100M1 likes355 downloads11mo agoHugging Face17jobs-git /HPLT2.0_cleanedThis is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to https://hplt-project.org/datasets/v2.0 The Cleaned variant of HPLT Datasets v2.0 This is the cleaned variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here. The original JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/HPLT2.0_cleaned.fill-maskn>1T1 likes329 downloads2y agoHugging Face18TiWu-Lab /HPLT-zhtabular100M<n<1B0 likes255 downloads1y agoHugging Face19HPLT /OpenLID-v3 Dataset Description OpenLID-v3 is an updated version of the OpenLID-v2 dataset (see the CHANGELOG.md). Repository: https://github.com/hplt-project/openlid Paper: OpenLID-v3: Improving the Precision of Closely Related Language Identification – An Experience Report Usage from datasets import load_dataset ds = load_dataset('HPLT/OpenLID-v3', split='train') Dataset Summary The OpenLID-v3 dataset covers 194 language varieties + not-a-language class… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/OpenLID-v3.texttext-classification100M<n<1B0 likes244 downloads4mo agoHugging Face20HPLT /HPLT3.0 This is a large-scale collection of web-crawled documents in 198 world languages, produced by the HPLT project. The source of the data is Internet Archive and Common Crawl. For a detailed description of this and previous releases by HPLT, please refer to our website. NB: the HPLT datasets are not hosted on HuggingFace! See download instructions below. HPLT release v3.0 In July 2025, the European HPLT initiative has completed a new release of its monolingual datasets, offering… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT3.0.fill-maskn>1T23 likes203 downloads3mo agoHugging Face21Eurolingua /HPLT3_DE_0.9_Quantile_Adult_Filteredtabular1M<n<10M1 likes202 downloads8mo agoHugging Face22HPLT /DocHPLT-v3-clean DocHPLT v3 cleaned Cleaned and re-formatted version of DocHPLTv3 ready for LLM training or document-level NMT training. texttranslation10M<n<100M1 likes197 downloads25d agoHugging Face23Finnish-NLP /HPLT_Finnish_fineweb_edu_predictedtabulartext-generation1M<n<10M0 likes189 downloads2y agoHugging Face24Ba2han /hplt-filteredtext10M<n<100M0 likes186 downloads2mo agoHugging Face25SlayerLab /hplt-v3-pl-cleaned HPLT v3 Polish — Cleaned & PII-Gated Oczyszczony, polski podzbiór korpusu web HPLT v3, przygotowany jako materiał pretreningowy. Autor / kurator zbioru: Arkadiusz Słota (SlayerLab). Wartość dodana względem surowego HPLT: wieloetapowa bramka PII (usuwanie numerów telefonów, identyfikatorów, adresów) z niezależną weryfikacją na pełnych danych + lekkie czyszczenie boilerplate. Wersja: v1.0 — floor (bins 8_5 + 8_6). Track B (bins 9_1 + 8_1, po deduplikacji względem bazy dynaword)… See the full description on the dataset page: https://huggingface.co/datasets/SlayerLab/hplt-v3-pl-cleaned.texttext-generation10M<n<100M0 likes185 downloads2mo agoHugging Face26tartuNLP /lumiopen-hpltv2-llama33-edu-annotation-ettabular100K<n<1M0 likes130 downloads1y agoHugging Face27Finnish-NLP /HPLT_1.2_fi_cleanedtabular1M<n<10M0 likes90 downloads3y agoHugging Face28HPLT /2508-datasets-evals HPLT 3.0: Details on Corpus Comparison Results Dataset Description This dataset contains fine-grained results from our HPLT 3.0 release evaluations comparing the new HPLT 3.0 corpora with the previous HPLT 2.0 version, FineWeb2, and MADLAD-400. We pretrain 2.2B Llama-style decoder models on 100B tokens for each selected language and evaluate them using HPLT-E, a multilingual evaluation framework for comprehensive multi-prompt k-shot evaluation across 124 tasks and 500+… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/2508-datasets-evals.tabular10K<n<100K0 likes82 downloads11mo agoHugging Face29nhagar /hplt-v1.2_urls Dataset Card for hplt-v1.2_urls This dataset provides the URLs and top-level domains associated with training records in HPLT v1.2. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it allows… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/hplt-v1.2_urls.text1B<n<10B0 likes77 downloads1y agoHugging Face30cpral /hplt3_pol_8_9_10_apt_tokenized_split0 likes71 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.