Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pdfqa /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Benchmark.documentquestion-answering5 likes7k downloads7mo agoHugging Face02Edinburgh-Claire /pdfQA-Benchmark pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. The dataset is organized to support: Raw document processing research Structured extraction pipelines Retrieval-augmented QA End-to-end document reasoning systems It preserves original documents alongside structured derivatives to enable reproducible evaluation across preprocessing strategies.… See the full description on the dataset page: https://huggingface.co/datasets/Edinburgh-Claire/pdfQA-Benchmark.documentquestion-answering0 likes3.4k downloads2mo agoHugging Face03pdfqa /pdfQA-Annotations pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs pdfQA is a structured benchmark collection for document-level question answering and PDF understanding research. This repository contains the pdfQA-Annotations dataset, which provides only the QA annotations and metadata for the pdfQA-Benchmark. It is intended for lightweight experimentation, modeling, and evaluation without requiring access to large document files. Relationship to the Full pdfQA… See the full description on the dataset page: https://huggingface.co/datasets/pdfqa/pdfQA-Annotations.documentquestion-answering3 likes445 downloads7mo agoHugging Face04zealwwww /gdp_pdf_aboard_aligned_5000gated GDP-PDF A-board-aligned 5K This gated dataset contains 5,000 Harbor-format professional document-reasoning tasks built from 5,000 unique public-sector PDFs across ten professional domains. Contents dataset/: Harbor task_XXXXX packages and index.json. sources/summary.json: conversion records and validation results. All 5,000 native packages and all 5,000 Harbor tasks passed the project validators. The task set contains 1,382 simple, 2,008 standard, and 1,610… See the full description on the dataset page: https://huggingface.co/datasets/zealwwww/gdp_pdf_aboard_aligned_5000.documentquestion-answering1K<n<10K2 likes164 downloads27d agoHugging Face05AingeruBeOr /RAG_legal_comparison_PDFs Dataset de Documentación Legal para Evaluación de Sistemas RAG Este dataset ha sido recopilado y estructurado en el marco de un Trabajo de Fin de Máster (TFM) enfocado en la comparativa de soluciones basadas en sistemas de Generación Aumentada por Recuperación (RAG) sobre documentación jurídica y administrativa. El objetivo principal es proporcionar un corpus realista y diverso de documentos en formato PDF para evaluar la capacidad de extracción de información, segmentación… See the full description on the dataset page: https://huggingface.co/datasets/AingeruBeOr/RAG_legal_comparison_PDFs.documenttext-retrievaln<1K0 likes86 downloads4mo agoHugging Face06CentificAIResearch /Healthcare.pdf Healthcare.pdf — representative release (v1.0) A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice. Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric. Tasks 25 Source documents… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf.documentquestion-answeringn<1K0 likes47 downloads11d agoHugging Face07ram-lexsi /curatorkit-testrun-PDF curatorkit-testrun-PDF Built using CuratorKIT — provenance-grounded curation and synthesis for LLM post-training. Method qa Backend litellm Model openai/Qwen/Qwen2.5-0.5B-Instruct Formats alpaca, sharegpt Artifact dataset Published 2026-09-01 06:27 UTC Usage from datasets import load_dataset ds = load_dataset("ram-lexsi/curatorkit-testrun-PDF", "alpaca") texttext-generationn<1K0 likes31 downloads1mo agoHugging Face08orgrctera /legalbenchrag_contractnli_retrieval_pdf legalbenchrag_contractnli_retrieval_pdf LegalBenchRAG contractnli retrieval dataset Field Value Benchmark legalbenchrag Sub-benchmark contractnli Type retrieval Items 977 Exported from Langfuse. textquestion-answeringn<1K0 likes26 downloads7mo agoHugging Face09orgrctera /legalbenchrag_cuad_retrieval_pdf legalbenchrag_cuad_retrieval_pdf LegalBenchRAG cuad retrieval dataset Field Value Benchmark legalbenchrag Sub-benchmark cuad Type retrieval Items 4042 Exported from Langfuse. textquestion-answering1K<n<10K0 likes25 downloads7mo agoHugging Face10Decre99 /Pdftextquestion-answeringn<1K0 likes18 downloads3y agoHugging Face11Roy229 /github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-financial-question-answering Financial Question Answering Dataset Summary Question-answer pairs extracted from financial documents and earnings reports. Dataset Structure Data fields: context, question, answer. Licensing Information This dataset is released under the MIT license (mit). question-answering0 likes15 downloads2mo agoHugging Face12UyopiQ /PDFtestdata Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/UyopiQ/PDFtestdata.textquestion-answeringn<1K0 likes13 downloads6mo agoHugging Face13Katya-Iukhn /optic_QA_pdf_russianimagequestion-answering1K<n<10K1 likes11 downloads2y agoHugging Face14gpahal /pdf-rag-embed-benchThis is a benchmark dataset for PDF RAG embedding systems. See gpahal/pdf-rag-embed-bench for more details. documentquestion-answeringn<1K0 likes9 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.