Team Ai
Datasetpublic

CentificAIResearch/Healthcare.pdf

Healthcare.pdf — representative release (v1.0) A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice. Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric. Tasks 25 Source documents… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf.

sourceHugging Facecc-by-4.0updated 12d agoView on Hugging Face
0likes47downloads
Dataset Card

Healthcare.pdf — representative release (v1.0)

A PDF-grounding benchmark for healthcare document work: 25 expert-authored tasks grounded in 25 real healthcare documents, covering 7 occupations across clinical and pharmacy practice.

Each task puts a practitioner in a realistic situation, gives them a real document, and asks a sequence of sub-questions that must be answered from that document. Answers are graded against a four-tier rubric.

Tasks25
Source documents25 (CDC Stacks, FDA drug labels)
Sub-questions152
Rubric criteria1,513
Occupations7
Document pages350

Loading

python
from datasets import load_dataset

ds = load_dataset("CentificAIResearch/Healthcare.pdf", split="test")
row = ds[0]

row["case_setup"]                     # the practitioner's situation
row["sub_questions"][0]["question"]   # first sub-question
row["sub_questions"][0]["reference_answer"]
row["document"]                       # the source PDF (decoded)
row["rubric"]                         # grading criteria, G1-G4
row["anchors"]                        # where in the document the answers live

The source PDFs are embedded in the Parquet, and also available as plain files under pdfs/.

Fields

FieldDescription
task_idpersona/occupation/task_id
persona, occupation, aspectrole and the workflow area the task belongs to
doc_id, document_name, document_file, doc_pagessource document identity
documentthe source PDF
case_setupthe clinical/pharmacy situation the practitioner is in
sub_questionslist of {q_no, question, reference_answer, task_types, anchor_ids, derivation, premises_grounded}
anchorslist of {id, element_type, title, locator_quote, page, answer_bearing} — the document locations the task is grounded in; sub_questions.anchor_ids and rubric.source_anchor_id resolve against these
rubriclist of {rubric_id, q_no, tier, grading_mode, kind, criterion, source_type, source_anchor_id, task_types}
n_subquestions, n_criteriacounts
difficulty_band, n_models_solved, n_models_all_criteriaobserved difficulty (see below)

Rubric tiers

  • —G1 Domain — healthcare-wide conformance (grounding, citation, no fabrication)
  • —G2 Persona — expectations for the persona
  • —G3 Occupation — expectations for the specific occupation
  • —G4 Task — the task's own content criteria, graded per sub-question

grading_mode is match (the answer must state this value), entailed (the answer must imply this step), or present (a behaviour/format requirement over the whole deliverable).

Task types

Sub-questions are tagged with a capability taxonomy (T1–T19) spanning foundational extraction (value extraction, layout navigation), structural reading (table parsing, conditional logic), advanced reasoning (abstention, quantitative reasoning, synthesis) and healthcare-specific work (clinical safety, care-transition extraction, eligibility matching). All 18 types present in the source pool are represented here.

Difficulty

difficulty_band records how many of six frontier models solved each task, under a two-grader consensus. A task is solved when every sub-question conveys the reference answer's key value or conclusion.

BandModels solvingTasks
unsolved0 / 66
hard1–2 / 65
mid3–4 / 66
easy5 / 64
ceiling6 / 64

n_models_all_criteria is a stricter reading — how many models passed every rubric criterion, not just conveyed the right answers. Across the release that is 8 of 150 model-task pairs versus 72 of 150 for solved. The gap is the point: most answers that look right still miss rubric criteria.

The hard end is deliberately over-sampled relative to the source pool, so this subset is harder than the full benchmark and its absolute scores should not be read as corpus-wide performance.

Source data

  • —CDC Stacks (14 documents) — US public-domain government publications
  • —FDA drug labels (11 documents) — US public-domain regulatory documents

Documents were selected for containing determinate, answer-bearing content that a practitioner would act on. Tasks and rubrics were authored against those documents and anchored to specific locations within them.

This release is the redistributable portion of a larger benchmark. A third track built on MIMIC-IV radiology reports is not included: PhysioNet credentialed access forbids redistribution.

Selection

The 25 were chosen from a 100-task redistributable pool to span difficulty, occupation and capability while preserving the pool's model ranking (Spearman ρ = 0.956, no inversions among separable model pairs). Rank fidelity was an explicit selection objective, so it is a design property rather than independent evidence. One task per source document; at most two tasks per workflow aspect.

Limitations

  • —Not a random sample. Deliberately over-weighted toward hard tasks; not an estimate of corpus-wide difficulty.
  • —Two personas only. Generalist clinicians (14) and pharmacists (11); the specialist track is excluded for licensing.
  • —Answer keys are public. Reference answers and rubrics ship with the tasks, so results are only meaningful for models that have not trained on this data. Treat it as a one-shot measurement and report the model's training cutoff.
  • —Expert review is partial. 6 of the 25 tasks were audited by practising clinicians as part of a broader purposive review; the remainder were not individually SME-reviewed.
  • —Grader-dependent. Difficulty metadata comes from LLM graders (two independent vendors, consensus required). Grader choice moves absolute numbers.

Licence

CC-BY-4.0. Source documents are US government works in the public domain; the tasks, rubrics and grounding annotations are original work released under CC-BY-4.0.

Citation

bibtex
@misc{healthcarepdf2026,
  title  = {Healthcare.pdf: a PDF-grounding benchmark for healthcare document work},
  author = {Centific AI Research},
  year   = {2026},
  url    = {https://huggingface.co/datasets/CentificAIResearch/Healthcare.pdf}
}