datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pdfQA-Benchmark
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
The dataset is organized to support:
Raw document processing research
Structured extraction pipelines
Retrieval-augmented QA
End-to-end document reasoning systems
It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.pdfQA-Benchmark
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
The dataset is organized to support:
Raw document processing research
Structured extraction pipelines
Retrieval-augmented QA
End-to-end document reasoning systems
It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/Edinburgh-Claire/pdfQA-Benchmark.pdfQA-Annotations
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research.
This repository contains the pdfQA-Annotations dataset, which provides only the QA annotations and metadata for the pdfQA-Benchmark.
It is intended for lightweight experimentation, modeling, and evaluation without requiring access to large document files.
Relationship to the Full pdfQA… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Annotations.gdp_pdf_aboard_aligned_5000
GDP-PDF A-board-aligned 5K
This gated dataset contains 5,000 Harbor-format professional document-reasoning tasks built from 5,000 unique public-sector PDFs across ten professional domains.
Contents
dataset/: Harbor task_XXXXX packages and index.json.
sources/summary.json: conversion records and validation results.
All 5,000 native packages and all 5,000 Harbor tasks passed the project validators. The task set contains 1,382 simple, 2,008 standard, and 1,610… See the full description on the dataset page: https://huggingface.co/datasets/zealwwww/gdp_pdf_aboard_aligned_5000.RAG_legal_comparison_PDFs
Dataset de Documentación Legal para Evaluación de Sistemas RAG
Este dataset ha sido recopilado y estructurado en el marco de un Trabajo de Fin de Máster (TFM) enfocado en la comparativa de soluciones basadas en sistemas de Generación Aumentada por Recuperación (RAG) sobre documentación jurídica y administrativa.
El objetivo principal es proporcionar un corpus realista y diverso de documentos en formato PDF para evaluar la capacidad de extracción de información, segmentación… See the full description on the dataset page: https://huggingface.co/datasets/AingeruBeOr/RAG_legal_comparison_PDFs.Healthcare.pdf
Healthcare.pdf — representative release (v1.0)
A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in
25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice.
Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of
sub-questions that must be answered from that document. Answers are graded against a four-tier rubric.
Tasks
25
Source documents… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf.curatorkit-testrun-PDF
curatorkit-testrun-PDF
Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training.
Method
qa
Backend
litellm
Model
openai/Qwen/Qwen2.5-0.5B-Instruct
Formats
alpaca, sharegpt
Artifact
dataset
Published
2026-09-01 06:27 UTC
Usage
from datasets import load_dataset
ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca")
legalbenchrag_contractnli_retrieval_pdf
legalbenchrag_contractnli_retrieval_pdf
LegalBenchRAG contractnli retrieval dataset
Field
Value
Benchmark
legalbenchrag
Sub-benchmark
contractnli
Type
retrieval
Items
977
Exported from Langfuse.
legalbenchrag_cuad_retrieval_pdf
legalbenchrag_cuad_retrieval_pdf
LegalBenchRAG cuad retrieval dataset
Field
Value
Benchmark
legalbenchrag
Sub-benchmark
cuad
Type
retrieval
Items
4042
Exported from Langfuse.
Pdfgithub_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-financial-question-answering
Financial Question Answering
Dataset Summary
Question-answer pairs extracted from financial documents and earnings reports.
Dataset Structure
Data fields: context, question, answer.
Licensing Information
This dataset is released under the MIT license (mit).
PDFtestdata
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/UyopiQ/PDFtestdata.optic_QA_pdf_russianpdf-rag-embed-benchThis is a benchmark dataset for PDF RAG embedding systems.
See gpahal/pdf-rag-embed-bench for more details.
