datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers.js-docspolicy-docs
Public Policy at Hugging Face
AI Policy at Hugging Face is a multidisciplinary and cross-organizational workstream. Instead of being part of a vertical communications or global affairs organization, our policy work is rooted in the expertise of our many researchers and developers, from Ethics and Society Regulars and legal team to machine learning engineers working on healthcare, art, and evaluations.
What we work on is informed by our Hugging Face community needs and experiences… See the full description on the dataset page: https://huggingface.co/datasets/huggingface/policy-docs.docs-imagesdiffusers-images-docsdolma3-6t-sample-100000-docs
dolma3-6t-sample-100000-docs
Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/).
Layout
HCAI-Lab/dolma3-6t-sample-100000-docs/
├── bin_summary.csv
├── sample_contract.json
└── worker_NNNN/
└── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker)
Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.pixmo-docs
PixMo-Docs
We now recommend using CoSyn-400k and CoSyn-point over these
datasets. They are improved versions with more images categories and an improved generation pipeline.
PixMo-Docs is a collection of synthetic question-answer pairs about various kinds of computer-generated images, including charts, tables, diagrams, and documents.
The data was created by using the Claude large language model to generate code that can be executed to render an image,
and using GPT-4o mini to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-docs.docsprotein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.argilla_sdk_docs_raw_unstructured
Dataset info
This dataset contains documentation chunks from repositories (ADD REPOS).
Postprocessing
After some inspection, some chunks contain text too short to be meaningful, so we decided to remove those by removing chunks whose number of tokens (computed
with the same tokenizer of the model to be used for the embeddings) is lower or equal to the 5%:
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("BAAI/bge-base-en-v1.5")
df =… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/argilla_sdk_docs_raw_unstructured.lk-news-docsengineering-docsdolma3-6t-sample-10000-docs
dolma3-6t-sample-10000-docs
Materialized stratified sample of 10K docs per bin (5.68M total docs, 10.5B tokens). Seed 42. This is the basis for the SOC-156 TrackStar gradient index.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_10000_docs
Renamed
2026-05-25… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs.elements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
lk-tourism-weekly-reports-docscbsl-annual-reports-docslk-dmc-weather-forecasts-docsdolma3-6t-sample-5000-docs
dolma3-6t-sample-5000-docs
Materialized stratified sample of 5K docs per bin (2.86M total docs, 5.3B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_5000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-5000-docs.lk-dmc-river-water-level-and-flood-warnings-docsdolma3-6t-sample-500-docs
dolma3-6t-sample-500-docs
Materialized stratified sample of 500 docs per bin (288K total docs, 539M tokens) drawn from the deduplicated 6T Dolma3 corpus. Seed 42. Matching companion bucket at hf://buckets/HCAI-Lab/dolma3-6t-sample-500-docs.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-500-docs.UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.dolma3-6t-sample-1000-docs
dolma3-6t-sample-1000-docs
Materialized stratified sample of 1K docs per bin (575K total docs, 1.08B tokens). Seed 42.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/dolma3_6T_sample_1000_docs
Renamed
2026-05-25
See docs/data_home/inventory.json for the full inventory… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-1000-docs.bhyt-legal-docsDocSynth300K
DocSynth300K is a large-scale and diverse document layout analysis pre-training dataset, which can largely boost model performance.
Data Download
Use following command to download dataset(about 113G):
from huggingface_hub import snapshot_download
# Download DocSynth300K
snapshot_download(repo_id="juliozhao/DocSynth300K", local_dir="./docsynth300k-hf", repo_type="dataset")
# If the download was disrupted and the file is not complete, you can resume the download… See the full description on the dataset page: https://huggingface.co/datasets/juliozhao/DocSynth300K.MAIR-Docs
MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
MAIR is a heterogeneous IR benchmark that comprises 126 information retrieval tasks across 6 domains, with annotated query-level instructions to clarify each retrieval task and relevance criteria.
This repository contains the document collections for MAIR, while the query data are available at https://huggingface.co/datasets/MAIR-Bench/MAIR-Queries.
Paper: https://arxiv.org/abs/2410.10127
Github:… See the full description on the dataset page: https://huggingface.co/datasets/MAIR-Bench/MAIR-Docs.huggingface-docsThis repo contains all the docs published on https://huggingface.co/docs.
The docs are generated with https://github.com/huggingface/doc-builder.
docs-gifs
FLIP documentation GIFs
The animated walkthroughs embedded in the FLIP user guides on
ReadTheDocs. They are recorded by Cypress from the
FLIP UI against a mocked backend (flip-ui/test/cypress/docs/), converted with ffmpeg, and published here by
.github/workflows/regenerate_docs_gifs.yml so the recordings never enter the git repository
(FLIP#1236).
Layout
One copy of every file, at an unversioned path:
Path
Contents
admin/<name>.gif
one per… See the full description on the dataset page: https://huggingface.co/datasets/aicentreflip/docs-gifs.n8n-docs-v2
n8n Docs
This repository hosts the documentation for n8n, an extendable workflow automation tool which enables you to connect anything to everything. The documentation is live at docs.n8n.io.
Previewing and building the documentation locally
Prerequisites
Python 3.8 or above
Pip
n8n recommends using a virtual environment when working with Python, such as venv.
Follow the recommended configuration and auto-complete guidance for the theme. This will help when… See the full description on the dataset page: https://huggingface.co/datasets/Jeice/n8n-docs-v2.text-code-galeras-code-generation-from-docstring-3k-dedupedlamini_docs
Dataset Card for "lamini_docs"
More Information needed
polars-docs
