BIDS
Datasets
All datasets matching “BIDS”arc-aphasia-bids
Aphasia Recovery Cohort (ARC)
Multimodal neuroimaging dataset for stroke-induced aphasia research.
Dataset Summary
The Aphasia Recovery Cohort (ARC) is a large-scale, longitudinal neuroimaging dataset containing multimodal MRI scans from 230 chronic stroke patients with aphasia. This HuggingFace-hosted version provides direct Python access to the BIDS-formatted data with embedded NIfTI files.
Metric
Count
Subjects
230
Sessions
902
T1-weighted scans
444… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/arc-aphasia-bids.medpmc-11m-dataset_jun24_baseline
MedPMC WebDataset
MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources.
This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.hbn-multimodal-bids-datasetM3LLM-data-v1.0.0
M3LLM Data
Data for M³LLM training and evaluation on biomedical instruction-following tasks derived from PubMed Central (PMC) articles. This repository is the versioned v1.0.0 data release.
Contents
Collection
Split
Records
Description
PMC-MI supervised instruction corpus
train
224,401
Six instruction formats after partitioning and release filtering
PMC-MI policy-refinement partition
train
10,355
Policy-refinement instances for Stage II after release… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data-v1.0.0.M3LLM-data
M3LLM-PMC Training Data
This dataset contains the training data for M3LLM (Medical Multimodal Large Language Model), comprising ~238K high-quality synthetic medical instruction-following samples.
Dataset Description
The data is generated from PubMed Central (PMC) medical literature through a comprehensive 5-stage synthetic data pipeline, covering six diverse medical visual question answering tasks.
Dataset Statistics
File
Samples
Task Type… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/M3LLM-data.medpmc-screening-dataset
MedPMC Initial Screening Training/Test Datasets
Overview
This dataset contains the annotations used for the initial screening stage of the MedPMC framework, which aims to identify clinically relevant medical images from biomedical literature.
The training and validation sets are automatically curated using GPT-4o. The test set is manually annotated.
For details on dataset construction, annotation guidelines, and the overall MedPMC pipeline, please refer to our… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-screening-dataset.
