Medical-Doctors-Last-Exam/mdle
MDLE: Case-Report Challenge Set — v1.3.0 357 questions · strict closed-book score < 6/10 for every evaluated solver model · no score of 6 included. This research challenge set is part of Medical Doctor's Last Exam, an independent clinical-reasoning benchmark maintained by the MDLE organization. It contains 357 English questions from 350 case-report articles in the March–July 2026 collection: 194 questions use images and 163 are text-only. The release bundles 235 original image… See the full description on the dataset page: https://huggingface.co/datasets/Medical-Doctors-Last-Exam/mdle.
MDLE: Case-Report Challenge Set — v1.3.0
357 questions · strict closed-book score < 6/10 for every evaluated solver model · no score of 6 included.
This research challenge set is part of Medical Doctor's Last Exam, an independent clinical-reasoning benchmark maintained by the MDLE organization. It contains 357 English questions from 350 case-report articles in the March–July 2026 collection: 194 questions use images and 163 are text-only. The release bundles 235 original image crops. Collection months are inherited source-group labels, not a guarantee that every article's publication date falls in that month.
Selection
Every included question satisfies both conditions:
gemini_closed_book_score < 6 AND opus_closed_book_score < 6The evaluated solvers are Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) and Claude Opus 5 (anthropic.claude-opus-5). “Every model” means these two existing closed-book evaluations. Source-assisted scores and judge models are not additional solver evaluations for this rule. Both scores must be present and valid. A score of 6 is excluded; there is no rounding, averaging, or less-than-or-equal comparison. The highest retained score for either model is 5.5/10.
There are no linked-Q1 exceptions. A Q1 is included only if its own two scores satisfy the strict rule. Compared with v1.2.1, this release retains 357 of 489 questions and removes 132: 113 former primary questions and all 19 supplemental Q1 context questions. The full retained/excluded item audit is in selection_changes.json.
Score and prompt versions
The selection threshold is 6, using a strict < comparison. Historical evaluation and correctness fields are preserved from v1.2.1, including their original passing score of 7. That historical grading threshold is distinct from this release's inclusion threshold, recorded in selection.threshold. Scores are out of 10.
The model scores were produced on the original v1.1.0 prompts. They were not rerun for v1.2.x or v1.3.0. This release preserves the v1.2.1 active prompts, reference answers, rubrics, original prompts and rewrite audits unchanged. For q02-fae0fd6e965a8cff, it retains the expert-reviewed injury-depth target and single-target rubric; the stored 0/10 scores still describe the earlier v1.1.0 two-target question.
The March source includes the Opus answer and detailed grading. The April–July Opus artifact contains only its scalar score, without the raw answer or detailed judge rationale. All retained rows carry historical wrong labels for both models; these are rubric-threshold labels, not claims that every statement in an answer is false.
Files and use
- data/test-00000-of-00001.parquet: the
defaultconfiguration andtestsplit, with native Hugging FaceList[Image]features and embedded image bytes. - data/challenge_2026_03_07.jsonl: the same 357 records as nested JSON Lines, without embedded image bytes.
assets/: the 235 required image crops; each row retains their repository-relative paths ininput_assets.- selection_manifest.json: current selection rule, counts, model identities, historical source metadata and per-item provenance.
- selection_changes.json: retained IDs and all 132 exclusions with the two scores.
- source_catalog.json: the 350 retained sources, citations, DOIs and article/PDF links.
- checksums.sha256: SHA-256 checksums for every release file except the checksum file itself.
- build_strict_release.py: reproducible builder, including the pinned parent revision and strict predicate.
from datasets import load_dataset
dataset = load_dataset(
"Medical-Doctors-Last-Exam/mdle",
revision="v1.3.0",
split="test",
)Each record includes the question, golden answer, reasoning, rubric, evidence, taxonomy, source identity, historical evaluations, correctness labels, current selection provenance, and the v1.2.x rewrite audit. The historical question_text_original is kept beside the active question_text.
Rebuild
Download the complete parent snapshot at revision de82a49729ee38215493d6e268468702406fa1bb to parent-v1.2.1/, including its image assets. After downloading this release, run from its root directory:
python -m pip install -r scripts/requirements.txt
python scripts/build_strict_release.py \
--source ../parent-v1.2.1 \
--assets ../parent-v1.2.1 \
--output ../rebuilt-v1.3.0 \
--repo-id Medical-Doctors-Last-Exam/mdleThe builder verifies parent and image checksums and refuses to overwrite the output directory. Record content and item membership are deterministic; the manifest creation timestamp and its checksum change on each rebuild.
Version history
- v1.3.0: strict two-model
<6/10subset; 357 questions; no linked-Q1 exceptions; no rewording or reevaluation. - v1.2.1: 489 questions; expert correction of the injury-depth target and rubric for
q02-fae0fd6e965a8cff. - v1.2.0: clearer active wording for all 489 questions, with original wording and audits retained.
- v1.1.0: 470 questions with both closed-book scores
<7/10, plus 19 linked Q1 context questions.
Parent dataset: Medical-Doctors-Last-Exam/mdle, revision de82a49729ee38215493d6e268468702406fa1bb. The parent JSONL SHA-256 is 2d5b170bad1c65f772e384a5abd7b06994e4152ae572fbfb3c72373b615b8e2d. Historical version tags remain available; pin a tag or full commit for reproducibility.
Limitations and rights
This is a deliberately difficult, model-selected set derived from case reports, not a representative sample of clinical questions or a licensing examination. Gemini 3.1 Pro Preview served as both an original solver and judge, including judging the Opus answers; selection may therefore reflect judge-specific bias. The historical scores are not current performance measurements or a general model ranking.
Original article PDFs are not redistributed. Source articles and images retain their authors' and publishers' rights; article licenses vary. The dataset retains source hashes and official links. Original questions and project-authored annotations are licensed under CC BY 4.0 where the project holds or has been granted the necessary rights. Project-owned code is licensed under Apache License 2.0. Article text, clinical images and other third-party material retain their original licenses and are excluded from these grants. The repository metadata uses license: other because the bundle has mixed source rights. See LICENSE.md. This benchmark is for research evaluation, not patient-care decisions.
Citation
Until the accompanying Medical Doctor's Last Exam paper is available, cite the dataset repository URL with version v1.3.0 and the resolved commit.
