datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-PDF-CC-2023-14
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.phaseShift_shell_result_pdf
Phase Resonance / IRS-DCE
Topological Dynamics & Artificial Cognitive Physics
Open Structural Record of Basis-Relative Reorganization in Transformer Representation Space
All pdf Creative Creative Commons Attribution No Derivatives 4.0 International
This repository provides comprehensive PDF research materials and Python scripts and something for mathematical proofs in the field of AI.
📢 Achievement: total 51k+ Downloads!… See the full description on the dataset page: https://huggingface.co/datasets/meta13sphere/phaseShift_shell_result_pdf.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.MINT-1T-PDF-CC-2023-23
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.multi-ocr-source-pdfs
MultiFinBen OCR source documents
📄 Paper · 💻 Code · 🌐 The Fin AI
Raw source documents behind the OCR tasks of MultiFinBen — MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application (arXiv:2506.14028). Formerly TheFinAI/OCR_Task (the old name redirects here).
This is not an evaluation set. For the curated OCR benchmarks use en-ocr, es-ocr, gr-ocr and jp-ocr.
Contents
Sources: U.S. SEC EDGAR filings (PDF/HTML) and… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/multi-ocr-source-pdfs.MINT-1T-PDF-CC-2023-50
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.example-pdfbiorXiv-pdf
BiorXiv Pdf
BiorXiv PDF dataset is a collection of PDF documents gathered from the BiorXiv website. This initiative aims to democratize artificial intelligence research by providing researchers with access to readily available training datasets. It is part of our broader effort to publish open access research papers as collective datasets.
BiorXiv is a renowned preprint publication in the field of biology and related disciplines. It is operated by Cold Spring Harbor Laboratory (CSHL)… See the full description on the dataset page: https://huggingface.co/datasets/laion/biorXiv-pdf.pdf_ocrpdf_1024medrXiv-pdf
MedrXiv Pdf
Introducing MedrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024. This resource aims to facilitate artificial intelligence research and the training of domain-specific scientific models.
As part of our efforts to democratise knowledge in the scientific domain, we have compiled this dataset. While most papers included have non-restrictive and open access licences, certain PDFs may have additional restrictions.
Researchers are encouraged to… See the full description on the dataset page: https://huggingface.co/datasets/laion/medrXiv-pdf.pdf-samplesocr-pdf-degraded
OCR-PDF-Degraded Dataset
Overview
This dataset contains synthetically degraded document images paired with their ground truth OCR text. It addresses a critical gap in OCR model training by providing realistic document degradations that simulate real-world conditions encountered in production environments.
Purpose
Most OCR models are trained on relatively clean, perfectly scanned documents. However, in real-world applications, especially in the military/defense… See the full description on the dataset page: https://huggingface.co/datasets/racineai/ocr-pdf-degraded.pdf_imagees-ocr-source-pdfs
MultiFinBen Spanish OCR source documents
📄 Paper · 💻 Code · 🌐 The Fin AI
Raw source documents behind the OCR tasks of MultiFinBen — MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application (arXiv:2506.14028). Formerly TheFinAI/MultiFinBen_OCR_Task (the old name redirects here).
This is not an evaluation set. For the curated OCR benchmarks use en-ocr, es-ocr, gr-ocr and jp-ocr.
Contents
Sources: Peruvian financial… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/es-ocr-source-pdfs.pdf_images_filtered
Image Dataset with Parquet Format
This dataset contains images with their IDs in parquet format for efficient loading.
Dataset Structure
Each configuration (language) contains:
image: PIL Image object
id: String identifier
Languages
Arabic (ar): 955 images
Bengali (bn): 932 images
German (de): 940 images
English (en): 932 images
Spanish (es): 951 images
French (fr): 947 images
Gujarati (gu): 949 images
Hindi (hi): 891 images
Italian (it): 1,005 images… See the full description on the dataset page: https://huggingface.co/datasets/v1v1d/pdf_images_filtered.OpenDoc-Pdf-Preview
OpenDoc-Pdf-Preview
OpenDoc-Pdf-Preview is a compact visual preview dataset containing 6,000 high-resolution document images extracted from PDFs. This dataset is designed for Image-to-Text tasks such as document OCR pretraining, layout understanding, and multimodal document analysis.
Dataset Summary
Modality: Image-to-Text
Content Type: PDF-based document previews
Number of Samples: 6,000
Language: English
Format: Parquet
Split: train only
Size: 606 MB
License: Apache… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDoc-Pdf-Preview.GS_PDF_repoHARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.fineweb-vi-pdf-layout
Fineweb VI PDF Layout (v2)
Vietnamese web pages (from HuggingFaceFW/fineweb-2) rendered to PDF, with per-page
layout annotations for VLM training.
Each row = 1 document:
Column
Type
Description
doc_name
string
Document ID (e.g. doc_000)
url
string
Source URL (metadata)
pdf
binary
Rendered document.pdf
layouts_merged
string
All per-page layout JSONs merged: {"num_pages": N, "pages": [...]}
pages
image list
Rendered page images (page_*.png)
source_html
string… See the full description on the dataset page: https://huggingface.co/datasets/dauvannam321/fineweb-vi-pdf-layout.pecan-pdfsKITAB_pdf_to_markdown_reviewed
KITAB_pdf_to_markdown_reviewed (Corrected KITAB-Bench PDF→Markdown)
Short description. A carefully reviewed and corrected version of the KITAB-Bench PDF-to-Markdown subset for Arabic document OCR evaluation. We fixed ground-truth errors (hallucinated text, missing page numbers, omissions of small-font text) and standardized formatting to provide a reliable benchmark for model comparison.
TL;DR
✅ Human-verified ground truth for Arabic PDF→Markdown
✅ Removes hallucinations and fills… See the full description on the dataset page: https://huggingface.co/datasets/Misraj/KITAB_pdf_to_markdown_reviewed.pdfvqa
Dataset Card for "pdfvqa"
More Information needed
chemrXiv-pdf
ChemrXiv Pdf
Introducing ChemrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024. This resource aims to facilitate artificial intelligence research and the training of domain-specific scientific models.
As part of our efforts to democratize knowledge in the scientific domain, we have compiled this dataset. While most papers included have non-restrictive and open access licenses, certain PDFs may have additional restrictions.
Researchers are encouraged… See the full description on the dataset page: https://huggingface.co/datasets/laion/chemrXiv-pdf.colpali_train_set_pdf_subset_sample_with_labelsfinance-pdf-vqaresumes-raw-pdf-for-ocrExtracted lists of pages from PDF resumes and the PDF texts.
Created using this code:
import io
import PIL.Image
from datasets import load_dataset
def render(pdf):
images = []
for page in pdf.pages:
buffer = io.BytesIO()
page.to_image(height=840).save(buffer)
images.append(PIL.Image.open(buffer))
return images
def extract_text(pdf):
return "\n".join(page.extract_text() for page in pdf.pages)
ds = load_dataset("d4rk3r/resumes-raw-pdf", split="train")… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/resumes-raw-pdf-for-ocr.fineweb-pdf-viekhmer-doc-markdown-pdf-kh3tables-pdf-20260823-pages
