datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SCPWiki-Cleaned-PDF-Archivessea-pdf-textpdf-parse-bench
PDF Parse Bench
Benchmark for evaluating how effectively PDF parsing solutions extract mathematical formulas and tables from documents.
We generate synthetic PDFs with diverse formatting scenarios, parse them with different parsers, and score the extracted content using LLM-as-a-Judge. This semantic evaluation approach substantially outperforms traditional metrics in agreement with human judgment.
Leaderboard (2026-Q1)
Results are based on two benchmark… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/pdf-parse-bench.pdf-parse-bench
PDF Parse Bench
Benchmark for evaluating how effectively PDF parsing solutions extract mathematical formulas and tables from documents.
We generate synthetic PDFs with diverse formatting scenarios, parse them with different parsers, and score the extracted content using LLM-as-a-Judge. This semantic evaluation approach substantially outperforms traditional metrics in agreement with human judgment.
Leaderboard (2026-Q1)
Results are based on two benchmark… See the full description on the dataset page: https://huggingface.co/datasets/ankitt6174/pdf-parse-bench.llm-jp-corpus-v4-ja_sip_comprehensive_pdf
llm-jp-corpus-v4 — ja_sip_comprehensive_pdf
Mirror of the ja/ja_sip_comprehensive_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_sip_comprehensive_pdf
Files: 156 × jsonl.gz (39.1 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_sip_comprehensive_pdf.jfk-pdfs-ocr
JFK OCRed files
Use colab: https://colab.research.google.com/drive/1xsXHYIaEU50tS0eTwuI5QdkGrbvjHQpW?usp=sharing to get more records.
SCPWiki-Cleaned-PDF-ArchivesswangchanLION-Curated-pdfaws-public-pdf-chunked-dataset
📚 AWS PDF Chunk Dataset
This dataset consists of chunked text extracted from all publicly available PDF documents on the Amazon Web Services (AWS) official website. The data includes user guides, technical documentation, and best practices about every AWS service, concept, and architecture.
It is designed to use in embedding generation, vector databases, and retrieval-augmented generation (RAG) systems.
📦 Dataset Structure
The dataset is provided as a .json file.… See the full description on the dataset page: https://huggingface.co/datasets/semihk1/aws-public-pdf-chunked-dataset.pdf-datasetPDF2TEX_raw_withgtpdf_filesterminal_filesystem_yahoo-finance_pdf-tools_huggingface_1990_settlement_a7c3f91epdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_aurora
Aurora Flow - Release Manifest
Candidate code name: aurora
Version: 2.1.0
Status: stable
License: MIT
Language: Python
Description: Real-time streaming data processing library
Maintainer: NovaTech Data Engineering
Last release: 2026-07-30
This dataset contains manifest.json (the release manifest) and validate.py
(the standardized validation script). Run python validate.py from this
directory to validate the release manifest.
pdf-3dsimulationdesign_compiler_pdf_datasetpdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_crimson
Crimson Stream - Release Manifest
Candidate code name: crimson
Version: 0.9.0
Status: beta
License: GPL-3.0
Language: Python
Description: Low-latency event streaming engine
Maintainer: NovaTech Data Engineering
Last release: 2026-08-02
This dataset contains manifest.json (the release manifest) and validate.py
(the standardized validation script). Run python validate.py from this
directory to validate the release manifest.
Coptic-PDF-Corpusimslp-raw-pdf-collection
IMSLP Raw PDF Collection
This dataset stores raw public-domain IMSLP PDF downloads collected by ReScore.
The raw PDF archive is append-only. Each Modal download launch publishes tar
shards under pdf_shards/ and matching JSONL manifests under manifests/.
Default collection id: imslp_pdf_collection.
The coordinator remains the source of claim/download state. Manifest rows link
each tar member back to the IMSLP id, source metadata, PDF checksum, page count,
Modal profile, launch id… See the full description on the dataset page: https://huggingface.co/datasets/cminst/imslp-raw-pdf-collection.filesystem_huggingface_yahoo_pdf_excel_pw_term_schol_3225_pr_a2ttai32pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_manifest_bluewave
BlueWave ETL - Release Manifest
Candidate code name: bluewave
Version: 1.4.0
Status: stable
License: Apache-2.0
Language: Python
Description: Batch ETL orchestration framework
Maintainer: NovaTech Data Engineering
Last release: 2026-06-15
This dataset contains manifest.json (the release manifest) and validate.py
(the standardized validation script). Run python validate.py from this
directory to validate the release manifest.
pdf-frame-dataset-1for-pdf-jsonlpdf-tools_huggingface_terminal_filesystem_7936_contract_register_t6zf5ppdf_to_quizz_mistralpdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_assessment
NovaTech Component Assessment Tracker
This dataset is the decision tracking repository for the vendor technology
assessment. After analysing the proposals and validating each candidate, the
analyst must upload the assessment report into this repository as:
assessment_report.json
assessment_report.csv
Acceptance policy:
A candidate is APPROVED when its validation passes AND its license is one of
MIT or Apache-2.0.
Otherwise it is REJECTED.
The final recommendation is the… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/pdf-tools_filesystem_github_terminal_huggingface_2118_x94fu9_assessment.pdftdtu_pdf_dataPDF-FORMSpdfjsonl
