datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
documentation-images
This dataset contains images used in the documentation of HuggingFace's libraries.
HF Team: Please make sure you optimize the assets before uploading them.
My favorite tool for this is https://tinypng.com/.
documentation-imagesdocument-haystack
Document Haystack Dataset
This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark”.
📑 Abstract Paper
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the processing of long documents remains under-explored, largely due to a lack of suitable benchmarks. To… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/document-haystack.wikitext_document_level
Wikitext Document Level
This is a modified version of https://huggingface.co/datasets/wikitext that returns Wiki pages instead of Wiki text line-by-line. The original readme is contained below.
Dataset Card for "wikitext"
Dataset Summary
The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified
Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/wikitext_document_level.documentation-imagesdocumentation-mediadocumentation-imagesCompany-document-dataset-v2
Company Documents v2
Generation complete: all 13 document types have completed export and upload checkpoints.
Synthetic, born-digital business documents rendered from four open sample databases, with exact gold
labels: 353,580 PDFs (404,514 pages) of 13 document types in
English and French, issued by 60 synthetic companies,
each with its own letterhead, numbering and wording. Successor of
CompanyDocuments (2,677 PDFs, 4 types).
Dataset overview
property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.documentation-imagesfoia-reading-room-documents
Foia Reading Room Documents
Documents from federal FOIA reading rooms and Inspector General report libraries: audits, inspections, investigative summaries and records released under the Freedom of Information Act.
Every document here was published by a US federal agency and is a work of the
United States government. Nothing has been altered: files are byte-identical to
what the agency posted, and the checksum in metadata.parquet is of the
original bytes.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/abigailhaddad/foia-reading-room-documents.documentation-imagesThis dataset contains images used in the documentation of HuggingFace's Optimum library.
document-review-data
Document Review Data
Private dataset for the Office/PDF title extraction review app and the current extractive title-training data package.
Current Title Extraction Dataset Surface
Canonical prefix:
datasets/title_extraction/
Effective datasets:
datasets/title_extraction/training/source4k_device_qwen_fp1000_v1/
datasets/title_extraction/evaluation/real_device_280_v1/
datasets/title_extraction/synthetic/controlled_synthetic_parse_v1/
The Dataset Viewer is… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-data.documentation-imagesDocumentVQAdocumentation-imagesJFK-Assassination-Records-2025-Documents-Releasedocumentation-imagesscotus-case-documents
Supreme Court of the United States: Case Documents
The document files that the Supreme Court's website lists under Case Documents in its footer, one config per collection, each row a file with the file itself, byte for byte, and its text. Nothing here is edited by hand, and no text is corrected, normalized or generated. The pipeline and its tests are in github.com/incrediblecrab/scotus-research-service-products, and this card is rendered from the collections' manifests in the… See the full description on the dataset page: https://huggingface.co/datasets/incrediblecrab/scotus-case-documents.Embrapa-ai-documents-markdownsynthetic-medical-document-recognition-benchmark
Synthetic Medical Document Recognition Benchmark
This dataset contains synthetic, English-language medical records rendered as
documents for evaluating automated data extraction and de-identification
systems. Each synthetic patient has a longitudinal FHIR R4 record and multiple
visual representations derived from that record.
Every rendered document is clearly marked as synthetic. This makes the dataset
suitable for manual testing, product demonstrations, and workflows that… See the full description on the dataset page: https://huggingface.co/datasets/morzel85/synthetic-medical-document-recognition-benchmark.form_understanding_in_noisy_scanned_documents_plus
Dataset Card for Form Understanding in Noisy Scanned Documents Plus
This is a FiftyOne dataset with 1026 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/form_understanding_in_noisy_scanned_documents_plus")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/form_understanding_in_noisy_scanned_documents_plus.AI2_Alphabot_2_stamp_document
AI2_Alphabot_2_stamp_document
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 987
Total Frames: 369702
FPS: 30
Dataset Size: 7.18 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb,
cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_stamp_document.jurisdb-legal-documents
JurisDB - Brazilian Legal Documents Dataset
Dataset Description
This dataset contains a comprehensive collection of Brazilian legal documents, including legislation (federal and state laws) and jurisprudence (court rulings and summaries from TSE, STJ, STF, and TNU).
Dataset Structure
.
├── legislacao_grifada_e_anotada_atualiz_em_01_01_2026/
│ ├── leis_estaduais/
│ ├── leis_federais/
│ └── ...
└── sumulas_tse_stj_stf_e_tnu_atualiz_01_01_2026_2/
├──… See the full description on the dataset page: https://huggingface.co/datasets/SherlockRamos/jurisdb-legal-documents.Agilex_Cobot_Magic_zip_up_the_document_bag
Agilex_Cobot_Magic_zip_up_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_zip_up_the_document_bag.document-haystack-10pages
Dataset Card for document-haystack-10pages
This is a FiftyOne dataset with 250 samples. It's the 10-page subset of the full dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/document-haystack-10pages")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/document-haystack-10pages.gardian-cigi-ai-documents-markdowndocumentation-imagesvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.document-review-source1k
New 1K title extraction corpus
Independent Task0178 expansion, parsed and machine annotated in Task0180, human-reviewed under Task0186. This is NOT the 1K subset sampled from the older 4K corpus. Private research use only; original document rights are not independently verified.
1000 unique source documents: Word, Excel, PDF and PPT each 250; English 500, simplified Chinese 400, traditional Chinese 100. All 1000 human reviews collected: 711 single-title, 194 multi-title, 54… See the full description on the dataset page: https://huggingface.co/datasets/mannycooper/document-review-source1k.ifpri-ai-documents-markdown
GAIA / GARDIAN-CIGI Agricultural Research Corpus
This dataset contains 21,726 agricultural research documents extracted from the GARDIAN repository and processed through the CIGI pipeline.
Dataset Overview
Property
Value
Total Documents
21,726
Total Size
623.27 MB
Total Tokens
85,359,442
Total Pages
0
Languages
25
Unique Keywords
7,127
Resource Types
20
Date Generated
2026-07-31 02:55:20
Language Distribution… See the full description on the dataset page: https://huggingface.co/datasets/CGIAR/ifpri-ai-documents-markdown.
