Yale-BIDS-Chen/M3LLM-data-v1.0.0
M3LLM Data Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release. Contents Collection Split Records Description PMC-MI supervised instruction corpus train 224,401 Six instruction formats after partitioning and release filtering PMC-MI policy-refinement partition train 10,355 Policy-refinement instances for Stage II after release… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data-v1.0.0.
M3LLM Data
Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release.
Contents
Initial PMC-MI assembly yielded 237,137 instruction instances. The source-article separation audit excluded 117 model-development records whose PMCIDs occurred in retained PMC-MI-Bench, yielding the final 237,020-instance study corpus. This corpus contains 226,464 supervised instructions and 10,556 policy-refinement instances, comprising 10,356 optimization and 200 validation instances.
The public release applies two additional provenance and licence filters. It excludes 43 records associated with 39 PMCIDs for which redistribution-licence evidence was not verified to the release threshold and 2,021 text-only records from the unavailable raw shard 008, whose PMCID and source-figure provenance could not be recovered. The resulting release contains 224,401 supervised records, 10,355 policy-refinement training records, 200 policy-refinement validation records and 298 benchmark records. metadata/training_exclusions.jsonl records each exclusion without redistributing its instruction content.
PMCID and source-figure provenance was recovered for 38,565 text-only source records: 38,366 by monotonic semantic alignment and 199 by exact context-question-answer alignment. After article-overlap and licence filtering, 38,516 text-only records are released. The alignment method and evidence are recorded per sample in metadata/puretext_provenance.jsonl.
Repository layout
data/pmc_mi/ PMC-MI supervised instruction corpus
data/policy_refinement/ PMC-MI policy-refinement train and validation splits
data/pmc_mi_bench/ evaluator-compatible benchmark JSON arrays
data/pmc_mi_bench_hf/ equivalent streamable JSONL files
images/pmc_mi_bench.tar.gz 588 images referenced by retained benchmark records
metadata/ sample manifests, article metadata and audits
scripts/ reproducible preparation and validation scriptsThe six benchmark filenames and their JSON-array format are preserved for direct use with the M3LLM-v2 evaluation code. Extract the benchmark images before evaluation:
mkdir -p M3LLM-v2/Evaluation/PMC-MI-Bench/images
tar -xzf images/pmc_mi_bench.tar.gz --strip-components=1 \
-C M3LLM-v2/Evaluation/PMC-MI-Bench/imagesIn text-only records, source_figure_id and the benchmark field images_puretext record provenance and are not model inputs. compound_image_id is retained as an internal compatibility field for the corresponding compound-figure identifier.
The policy-refinement files retain EasyR1-compatible fields. ground_truth_selection (and compatibility alias answer) stores reference subimage identifiers, ground_truth_response stores the structured reference trajectory and ground_truth_answer_text stores the refined reference answer.
Loading the data
from datasets import load_dataset
pmc_mi = load_dataset("KerwinFu/M3LLM-data-v1.0.0", "pmc_mi_multi_subimage_vqa")
policy = load_dataset("KerwinFu/M3LLM-data-v1.0.0", "pmc_mi_policy_refinement")
benchmark = load_dataset("KerwinFu/M3LLM-data-v1.0.0", "pmc_mi_bench_multi_subimage_vqa")Visual PMC-MI and policy-refinement records reference source-image filenames from the MedPMC-11M image repository. The 588 images referenced by PMC-MI-Bench are included here.
Provenance and licences
Every released sample has a PMCID and article-level licence metadata. The combined evidence comprises 234,288 sample records resolved from the frozen PMC AWS inventory, 767 from official PMC JATS and 199 from the labelled earlier PMC Open Access audit. The retained article licences are CC BY, CC BY-NC, CC BY-NC-SA, CC BY-SA or CC0.
The released training data and PMC-MI-Bench are disjoint at the PMCID and stored compound-figure-identifier levels. The complete sample mapping, article-version evidence, exclusions and overlap audit are provided under metadata/.
To the extent that the authors hold the relevant rights, project-generated questions, answer options, reference answers, structured reference trajectories, task labels and dataset organisation are licensed under CC BY-NC-SA 4.0, consistent with the MedPMC release. PMC-derived images, captions and article text remain governed by their per-source licences; accordingly, this repository uses a mixed-licence declaration.
See DATA_LICENCE.md for the licence scope and reuse guidance. Preparation counts and checks are recorded in metadata/preparation_summary.json, metadata/licence_coverage_summary.json, metadata/overlap_audit.json and metadata/file_checksums.sha256.
Associated resources
- Code: https://github.com/KerwinFuyihang/M3LLM-v2/tree/v1.0.2
- Source-image repository: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-datasetjun24baseline
