datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter.
The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object).
med_mcqaFrom "MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering"
(Pal et al.), MedMCQA is a "multiple-choice question answering (MCQA) dataset designed to address
real-world medical entrance exam questions." The dataset "...has more than 194k high-quality AIIMS & NEET PG
entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average
token length of 12.77 and high topical diversity."
The following is an example from… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/med_mcqa.copycolors_mcqa
Synthetic copycolors_mcqa (4 answer choices)
This dataset is a synthetic extension of
mib-bench/copycolors_mcqa,
restricted to the 4-choice setting used in this repository.
It keeps only these counterfactual families:
answerPosition_counterfactual
randomLetter_counterfactual
answerPosition_randomLetter_counterfactual
The export uses a single train split. Each row contains one base prompt and one source row
for each of the three counterfactual types, so the dataset is balanced… See the full description on the dataset page: https://huggingface.co/datasets/jchang153/copycolors_mcqa.Nemotron-RL-knowledge-web_search-mcqa
Dataset Description:
The Nemotron-RL-knowledge-web_search-mcqa is a multi-domain synthetic dataset designed to improve science and general reasoning in large language models (LLMs). It is a filtered subset of the OpenScienceReasoning-2 dataset and contains multiple-choice question–answer pairs spanning diverse domains: physics, biology, mathematics, humanities, computer science, engineering, chemistry, and others.
This dataset is released as part of NVIDIA NeMo Gym, a framework… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-web_search-mcqa.drtulu_v2_nemotron_web_search_mcqamath_mcqa
MATH-MCQA
A multiple choice adaptation of the MATH dataset containing 12,498 competition-level mathematics problems.
Key Statistics
Metric
Value
Total Examples
12,498
Train Split
7,498
Test Split
5,000
Categories
7 (Algebra, Intermediate Algebra, Prealgebra, Geometry, Number Theory, Counting & Probability, Precalculus)
Difficulty Levels
5 (Level 1 = Easiest, Level 5 = Hardest/Competition-level)
Format
4-option multiple choice (1 correct answer + 3… See the full description on the dataset page: https://huggingface.co/datasets/stellaathena/math_mcqa.med_mcqaFrom "MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering"
(Pal et al.), MedMCQA is a "multiple-choice question answering (MCQA) dataset designed to address
real-world medical entrance exam questions." The dataset "...has more than 194k high-quality AIIMS & NEET PG
entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average
token length of 12.77 and high topical diversity."
The following is an example from… See the full description on the dataset page: https://huggingface.co/datasets/ap878/med_mcqa.MRI-MCQA
MRI-MCQA
Dataset Description
MRI-MCQA is a benchmark composed by multiple-choice questions related to Magnetic Resonance Imaging (MRI). We use this dataset to evaluate the level of knowledge of various LLMs about the MRI field.
Curated by: Oscar Molina Sedano
Language(s) (NLP): English
License
This dataset is licensed under CC-BY-NC 4.0.
Disclaimer
Courtesy of Allen D. Elster… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MRI-MCQA.copycolors_mcqaThis dataset consists of a formatted version of the Memory Colors dataset (Norland et al. 2021), formatted for n-way multiple choice where n is in [2,11].
arc-synth-mcqa
ARC-Style Synthetic Science MCQs (teacher: Qwen3.8-27B)
Synthetic multiple-choice science questions generated for ARC-Challenge
fine-tuning, released for reproducibility of the companion model. Every file
that was used in training is included, along with the full audit trail.
Files
file
rows
what it is
clean_all.jsonl
6,861
v2 pool: generated, blind-label-verified, deduped, ARC-form-gated
clean_std.jsonl
4,630
non-negation subset of the above… See the full description on the dataset page: https://huggingface.co/datasets/minjujeon/arc-synth-mcqa.MNLP_M2_mcqa_dataset
Dataset Card for SCP-116K
Recent Updates
We have made significant updates to the dataset, which are summarized below:
Expansion with Mathematics Data:Added over 150,000 new math-related problem-solution pairs, bringing the total number of examples to 274,166. Despite this substantial expansion, we have retained the original dataset name (SCP-116K) to maintain continuity and avoid disruption for users who have already integrated the dataset into their workflows.
Updated… See the full description on the dataset page: https://huggingface.co/datasets/asazheng/MNLP_M2_mcqa_dataset.Medical-MCQA
NLI-Filtered Medical MCQs
This repository contains 650 evidence-grounded medical multiple-choice questions in JSONL format. The dataset was prepared from a larger question bank generated with a retrieval-augmented generation pipeline that retrieves supporting evidence from PubMed and textbook-style medical sources, then verifies each item before logging it.
Dataset Summary
Total rows: 650
File format: JSON Lines (dataset.jsonl)
Primary language: English
Optional… See the full description on the dataset page: https://huggingface.co/datasets/iqbalr/Medical-MCQA.MNLP_M2_mcqa_datasetMNLP_M2_mcqa_datasetnemotron-nano-rl-mcqa-19k
Nemotron Nano RL MCQA 19K
Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format.
Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.MNLP_M2_mcqa_dataMCQA-ExpMNLP_M3_mcqa_datasetSynthetic_Dataset_For_MCQAexample-thai-mcqasafety_aegis_mcqa
Safety aegis-to-MCQ transform
NVIDIA Nemotron-Content-Safety (aegis_v2) reformatted into SafetyBench-style MCQ (category/binary/safe-action), gold grounded in human prompt_label/violated_categories. Produced by data/prepare_aegis_mcqa.py (seed=42); hard_aegis = decision-boundary subset for the CLARITY/RLVR stage. Used to train the CS-552 safety_model.
UEH-MCQAmcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer (correct choice)', 'max_choices (number of answer options)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian_manual.MNLP_M2_mcqa_datasetmnlp_mcqa_evals_factualMNLP_M3_mcqa_datasetMNLP_M3_mcqa_datasetmcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices (number of answer options)', 'answer (correct choice)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian.mcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices', 'answer'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
mcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer', 'max_choices'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
