CentificAIResearch/Healthcare.pdf
Healthcare.pdf — representative release (v1.0) A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice. Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric. Tasks 25 Source documents… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf.
Healthcare.pdf — representative release (v1.0)
A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice.
Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric.
Loading
from datasets import load_dataset
ds = load_dataset("CentificAIResearch/Healthcare.pdf", split="test")
row = ds[0]
row["case_setup"] # the practitioner's situation
row["sub_questions"][0]["question"] # first sub-question
row["sub_questions"][0]["reference_answer"]
row["document"] # the source PDF (decoded)
row["rubric"] # grading criteria, G1-G4
row["anchors"] # where in the document the answers liveThe source PDFs are embedded in the Parquet, and also available as plain files under pdfs/.
Fields
Rubric tiers
- G1 Domain — healthcare-wide conformance (grounding, citation, no fabrication)
- G2 Persona — expectations for the persona
- G3 Occupation — expectations for the specific occupation
- G4 Task — the task's own content criteria, graded per sub-question
grading_mode is match (the answer must state this value), entailed (the answer must imply this step), or present (a behaviour/format requirement over the whole deliverable).
Task types
Sub-questions are tagged with a capability taxonomy (T1–T19) spanning foundational extraction (value extraction, layout navigation), structural reading (table parsing, conditional logic), advanced reasoning (abstention, quantitative reasoning, synthesis) and healthcare-specific work (clinical safety, care-transition extraction, eligibility matching). All 18 types present in the source pool are represented here.
Difficulty
difficulty_band records how many of six frontier models solved each task, under a two-grader consensus. A task is solved when every sub-question conveys the reference answer's key value or conclusion.
n_models_all_criteria is a stricter reading — how many models passed every rubric criterion, not just conveyed the right answers. Across the release that is 8 of 150 model-task pairs versus 72 of 150 for solved. The gap is the point: most answers that look right still miss rubric criteria.
The hard end is deliberately over-sampled relative to the source pool, so this subset is harder than the full benchmark and its absolute scores should not be read as corpus-wide performance.
Source data
- CDC Stacks (14 documents) — US public-domain government publications
- FDA drug labels (11 documents) — US public-domain regulatory documents
Documents were selected for containing determinate, answer-bearing content that a practitioner would act on. Tasks and rubrics were authored against those documents and anchored to specific locations within them.
This release is the redistributable portion of a larger benchmark. A third track built on MIMIC-IV radiology reports is not included: PhysioNet credentialed access forbids redistribution.
Selection
The 25 were chosen from a 100-task redistributable pool to span difficulty, occupation and capability while preserving the pool's model ranking (Spearman ρ = 0.956, no inversions among separable model pairs). Rank fidelity was an explicit selection objective, so it is a design property rather than independent evidence. One task per source document; at most two tasks per workflow aspect.
Limitations
- Not a random sample. Deliberately over-weighted toward hard tasks; not an estimate of corpus-wide difficulty.
- Two personas only. Generalist clinicians (14) and pharmacists (11); the specialist track is excluded for licensing.
- Answer keys are public. Reference answers and rubrics ship with the tasks, so results are only meaningful for models that have not trained on this data. Treat it as a one-shot measurement and report the model's training cutoff.
- Expert review is partial. 6 of the 25 tasks were audited by practising clinicians as part of a broader purposive review; the remainder were not individually SME-reviewed.
- Grader-dependent. Difficulty metadata comes from LLM graders (two independent vendors, consensus required). Grader choice moves absolute numbers.
Licence
CC-BY-4.0. Source documents are US government works in the public domain; the tasks, rubrics and grounding annotations are original work released under CC-BY-4.0.
Citation
@misc{healthcarepdf2026,
title = {Healthcare.pdf: a PDF-grounding benchmark for healthcare document work},
author = {Centific AI Research},
year = {2026},
url = {https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf}
}