Team Ai
Datasetpublic

Shivi1982/clin-auth-bench

ClinAuthBench ClinAuthBench is a synthetic inpatient health authorization benchmark. V1 focuses on adult inpatient psychiatric authorization over dense 72-hour chart packets. Each record contains a synthetic multi-form chart packet and structured gold labels for continued-stay reasoning, lower-level-of-care readiness, evidence grounding, risk reconciliation, and unsupported-claim avoidance. Links GitHub (evaluation code, baselines, generators):… See the full description on the dataset page: https://huggingface.co/datasets/Shivi1982/clin-auth-bench.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes45downloads
Dataset Card

ClinAuthBench

ClinAuthBench is a synthetic inpatient health authorization benchmark. V1 focuses on adult inpatient psychiatric authorization over dense 72-hour chart packets.

Each record contains a synthetic multi-form chart packet and structured gold labels for continued-stay reasoning, lower-level-of-care readiness, evidence grounding, risk reconciliation, and unsupported-claim avoidance.

Links

  • —GitHub (evaluation code, baselines, generators): https://github.com/Shivi1982/clin-auth-bench
  • —Paper / technical report: (to be added — arXiv preprint forthcoming)

Dataset Contents

  • —180 synthetic 72-hour chart packets.
  • —108 continued-stay cases.
  • —72 safe or lower-level-of-care-ready cases.
  • —22 contradiction/conflicting-documentation cases.
  • —120 cases using rule_based_v1 MDP trajectories.
  • —60 cases using probabilistic_v1 MDP trajectories.
  • —Cases 121-180 include probabilistic transition traces.

The release file is:

text
data/synthetic_bh_cases_v1_mdp_180.json

Loading Example

Install the Hugging Face datasets library and load the test split:

from datasets import load_dataset

dataset = load_dataset("Shivi1982/clin-auth-bench", split="test")

print(dataset[0]["id"]) print(dataset[0]["metadata"]["gold"]["safeforlloc"]) print(dataset[0]["content"][:500])

Record Schema

Each record has four top-level fields:

  • —id: stable case identifier.
  • —title: short synthetic case title.
  • —content: model-facing synthetic chart packet.
  • —metadata: benchmark metadata, generation metadata, QA checks, and gold labels.

Important nested fields include:

  • —metadata.gold.safe_for_lloc
  • —metadata.gold.expected_los_recommendation
  • —metadata.gold.lower_level_of_care_barriers
  • —metadata.gold.evidence_anchors
  • —metadata.gold.do_not_claim
  • —metadata.documentation_challenge
  • —metadata.documentation_challenge_tags
  • —metadata.mdp_model
  • —metadata.mdp_trajectory

See docs/schema.md for a fuller schema summary.

Distinctive Features

ClinAuthBench is designed to evaluate evidence discipline, not just answer accuracy.

  • —Explicit negative constraints: metadata.gold.do_not_claim identifies tempting unsupported conclusions the model should avoid.
  • —MDP trajectory metadata: case packets are generated from hidden synthetic state transitions over a 72-hour authorization window.
  • —Documentation-challenge taxonomy: v1 includes current-vs-historical risk, contradiction, lower-level-of-care barrier reasoning, and missing/invalid/stale structured evidence.
  • —Dense multi-form packets: cases require synthesis across nursing, psychiatry, social work, medication, group, rating-scale, and discharge-planning documentation.

Intended Uses

This dataset is intended for research and evaluation of:

  • —Clinical information extraction from dense inpatient documentation.
  • —Continued-stay authorization summarization.
  • —Source-grounded evidence citation.
  • —Current vs historical suicide-risk reconciliation.
  • —Contradiction detection across chart forms.
  • —Evaluation of hallucinated claims, fabricated evidence, invalid scores, and unsupported level-of-care recommendations.

ClinAuthBench is not intended to train or evaluate direct clinical care decisions.

Privacy And Affiliation Notice

ClinAuthBench is fully synthetic. It is generated from structured synthetic case specifications, Markov Decision Process-style synthetic trajectories, synthetic timelines, and controlled documentation templates.

It is not derived from, affiliated with, endorsed by, or representative of any healthcare provider, payer, employer, customer, EHR vendor, or real patient population.

The dataset does not contain real patient records, copied chart text, PHI, real facility names, real staff names, real patient names, MRNs, phone numbers, addresses, or source-system identifiers.

All chart dates are synthetic timeline artifacts and do not correspond to real patient encounters or real facility operations.

QA Summary

Pre-release QA results:

  • —Deterministic QA: PASS=180, REVIEW=0, FAIL=0.
  • —Claude PHI audit, random 30 full-content case sample: PASS=30, FAIL=0.

The latest local audit artifact used for this summary was audit_results_20260603T043315.json. Audit result JSON files are intentionally not included in this Hugging Face release package.

Limitations

  • —V1 contains 180 cases and is evaluation-scale, not training-from-scratch scale.
  • —V1 focuses on adult inpatient psychiatric authorization only.
  • —It does not include pediatric, geriatric, ED-only, PHP/IOP, outpatient, or residential-only cases.
  • —Synthetic notes are cleaner and more uniform than real EHR documentation.
  • —Each case uses a fixed 72-hour documentation window.
  • —Authorization criteria are generalized and are not tied to any payer, provider, facility, EHR, contract, or proprietary utilization-management policy.
  • —MDP trajectories are synthetic generation metadata, not observed real-world clinical transitions.

License

The dataset is released under the Creative Commons Attribution 4.0 International license (CC BY 4.0). Users may share and adapt the dataset with attribution.

Citation

bibtex
@misc{bhatia2026clinauthbench,
  author       = {Bhatia, Shivi},
  title        = {ClinAuthBench v1: A Synthetic Benchmark for Inpatient Health Authorization Reasoning},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/Shivi1982/clin-auth-bench}}
}

A formal arXiv citation will be added when the technical report is posted.