Team Ai
Datasetpublic

vidulpanickan/TinyEHR

TinyEHR v0.2.0 | GitHub | Website | PyPI A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use. This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.

sourceHugging Faceodblupdated 6mo agoView on Hugging Face
3likes563downloads
Dataset Card

TinyEHR

v0.2.0 | GitHub | Website | PyPI

A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use.

This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care.

This dataset is derived from real EHR data from Beth Israel Deaconess Medical Center (BIDMC) in Boston, US. The data has been de-identified, meaning it has been stripped of any information that could identify the patient such as names, medical record numbers, and addresses to protect patient privacy. This dataset contains no protected health information (PHI).

StatValue
Patients100
Hospital admissions275
ICU stays140
Clinical notes4,580
Gender43 F / 57 M
Date range2011 - 2022
Tables (MIMIC)33
Tables (OMOP)32

Browse the dataset: explore 30+ tables, column definitions, and relationships across MIMIC-IV and OMOP formats.

![TinyEHR Schema Explorer](https://tinyehr.org)

AI assisted SQL: ask your queries in plain English.

![TinyEHR AI SQL](https://tinyehr.org)

Quick Start

python
from datasets import load_dataset

patients = load_dataset("vidulpanickan/TinyEHR", "mimic_patients")
admissions = load_dataset("vidulpanickan/TinyEHR", "mimic_admissions")
notes = load_dataset("vidulpanickan/TinyEHR", "mimic_noteevents")

Also available as a Python package: pip install tinyehr (PyPI)

What does the data look like?

patients (subject_id = patient ID, anchor_age = age at anchor year, dod = date of death):

json
{
  "subject_id": 10014729,
  "gender": "F",
  "anchor_age": 21,
  "anchor_year": 2013,
  "anchor_year_group": "2011 - 2013",
  "dod": null
}

noteevents (hadm_id = hospital admission ID, note_type = type of clinical note):

json
{
  "note_id": "10014729-DS-0001",
  "subject_id": 10014729,
  "hadm_id": 23300884,
  "note_type": "Discharge summary",
  "chartdate": "2013-03-19",
  "text": "Admission Date: 2013-03-19  Discharge Date: 2013-03-28\n\nDOB: 1992  Sex: F\n\nService: VSURG → CSURG\n\nAttending: Dr. Katriel Silvane\n\nALLERGIES: NKDA\n\nCC: Postop wound infection s/p thoracotomy..."
}

There are 30+ tables covering admissions, diagnoses, lab results, medications, procedures, vitals, clinical notes, and more. Explore all tables at tinyehr.org.

Two Formats

FormatTablesRowsBest for
tinyehr_mimic_format33~1.4MLearning how hospital data works
tinyehr_omop_format32~472KBuilding tools that work across health systems

MIMIC-IV format follows the original MIMIC-IV schema. If you're new to EHR data, start here.

OMOP CDM v5.3.1 format reorganizes the same data into a universal schema where diagnoses, labs, and medications are mapped to standardized medical vocabularies.

Full details: ABOUT_THE_DATA.md

Usage

  • —Build and test AI agents that query, reason over, and navigate real hospital data
  • —Prototype clinical NLP and text-to-SQL systems against realistic clinical notes and multi-table schemas
  • —Learn how EHR data is structured across MIMIC-IV and OMOP formats

Known Limitations

  • —100 patients only: this is a learning and prototyping dataset, not statistically representative of any population
  • —Clinical notes are generated and not validated: the notes were generated using Anthropic's Claude Opus 4.6, grounded in each patient's structured data during their hospital visit. They have not been validated by clinicians and may contain hallucinated or inaccurate clinical details (e.g., incorrect ages, fabricated findings, inconsistent timelines). They should not be treated as clinically accurate
  • —Single institution: all data comes from one US academic medical center (Beth Israel Deaconess Medical Center in Boston), so demographics and clinical patterns reflect this specific patient population
  • —OMOP vocabulary subset: the OMOP format uses a subset of the full OHDSI Athena vocabulary, limited to the concepts needed for these 100 patients

Roadmap

  • —Synthetic clinical notes authored by clinicians (currently generated by LLM)
  • —Additional data modalities including medical imaging (X-ray, CT scan)

Citation

If you use TinyEHR in your work, please cite:

bibtex
@misc{tinyehr2026,
  title={TinyEHR: A 100 Patient Electronic Health Records Dataset for Learning and Prototyping Agentic AI},
  author={Vidul Ayakulangara Panickan},
  year={2026},
  url={https://github.com/vidulpanickan/TinyEHR}
}

Source Citations

  1. 1.Johnson, A., Bulgarelli, L., Pollard, T., Horng, S., Celi, L. A., & Mark, R. (2023). MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 10(1), 1. https://doi.org/10.1038/s41597-022-01899-x
  1. 1.Johnson, A., Bulgarelli, L., Pollard, T., Horng, S., Celi, L. A., & Mark, R. (2023). MIMIC-IV Clinical Database Demo (version 2.2). PhysioNet. https://doi.org/10.13026/dp1f-ex47
  1. 1.Kallfelz, M., Tsvetkova, A., Pollard, T., Kwong, M., Lipori, G., Huser, V., Osborn, J., Hao, S., & Williams, A. (2021). MIMIC-IV Demo Data in the OMOP Common Data Model (version 0.9). PhysioNet. https://doi.org/10.13026/p1f5-7x35

License

ODbL-1.0 (Open Data Commons Open Database License). Free to use, share, and modify. Redistributed versions must use the same license.