datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter.
The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object).
Nemotron-RL-knowledge-mcqa
Dataset Description:
The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.med_mcqaFrom "MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering"
(Pal et al.), MedMCQA is a "multiple-choice question answering (MCQA) dataset designed to address
real-world medical entrance exam questions." The dataset "...has more than 194k high-quality AIIMS & NEET PG
entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average
token length of 12.77 and high topical diversity."
The following is an example from… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/med_mcqa.driving_mcqa
DrivingExamMCQA
The DrivingExamMCQA dataset is a Multiple-Choice Question Answering (MCQA) collection based on real driving exam questions. It supports multilingual assessment across three languages: Arabic (ar), French (fr), and English (en) (with translations).
Overview
Each language includes two modalities:
Image-supported questions (_img splits):
Questions paired with an image (e.g., road signs, traffic scenarios).
Text-only questions (_text splits):
Standard… See the full description on the dataset page: https://huggingface.co/datasets/ESmike/driving_mcqa.science-mcqa-training-pool
Science multiple-choice training pool
Public multiple-choice science questions from three datasets, read at the pinned revisions named
below and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 182035 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
question
the question text, as its source publishes it
options
the answer options, as… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/science-mcqa-training-pool.NLU-Belebele-MCQAMNLP_MCQA_datasetThis MCQA dataset (of only single answer) contains a mixture of train, validation and test from this datasets (test and validation are only used for testing not for training):
mmlu auxiliary train Only the stem subset is used
mmlu Only the stem subset is used
mmlu 10 choices auxiliary train stem
ai2_arc
ScienceQA
math_qa
openbook_qa
sciq
medmcqa A 32,000 random subset (seed 42)
spider_mcqa_v0.2_full
Spider-MCQA
Converted Spider Text-to-SQL (Paper: Yu et al., 2018; HF Dataset) test set into multiple-choice.
The dataset contains 1,034 examples.
Dataset Fields
Each JSON record contains:
query: the schema and natural-language question prompt.
gold_answer: the correct SQL answer.
options: four SQL answer options, including the gold answer and three generated distractors.
correct_option_index: the index of the correct answer in options.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notpaulmartin/spider_mcqa_v0.2_full.wmdp_bio_robust_mcqaMNLP_M3_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets (and the test and validation splits are only used for testing not for training):
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M3_mcqa_dataset.MNLP_M3_mcqa_datasetAfri-MCQA
Afri-MCQA: Multimodal Cultural Question Answering for African Languages
Paper
Overview
Afri-MCQA is the first multilingual cultural question-answering benchmark covering 8k Q&A pairs across 16 African languages from 13 countries. The benchmark offers parallel English-African language Q&A pairs across text and speech modalities, entirely created by native speakers.
Supported Tasks
Visual Question Answering (VQA): Multiple-choice and open-ended QA… See the full description on the dataset page: https://huggingface.co/datasets/Atnafu/Afri-MCQA.MNLP_M3_mcqa_datasetMNLP_M3_mcqa_datasetmcqa_calibration_datasetcopycolors_mcqa
Synthetic copycolors_mcqa (4 answer choices)
This dataset is a synthetic extension of
mib-bench/copycolors_mcqa,
restricted to the 4-choice setting used in this repository.
It keeps only these counterfactual families:
answerPosition_counterfactual
randomLetter_counterfactual
answerPosition_randomLetter_counterfactual
The export uses a single train split. Each row contains one base prompt and one source row
for each of the three counterfactual types, so the dataset is balanced… See the full description on the dataset page: https://huggingface.co/datasets/jchang153/copycolors_mcqa.MNLP_M2_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets:
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the stem subsets
mmlu we only use the stem subsets
ai2_arc
ScienceQA
math_qa… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M2_mcqa_dataset.MNLP_M3_mcqa_dataset_openbookqa_cotNemotron-RL-knowledge-web_search-mcqa
Dataset Description:
The Nemotron-RL-knowledge-web_search-mcqa is a multi-domain synthetic dataset designed to improve science and general reasoning in large language models (LLMs). It is a filtered subset of the OpenScienceReasoning-2 dataset and contains multiple-choice question–answer pairs spanning diverse domains: physics, biology, mathematics, humanities, computer science, engineering, chemistry, and others.
This dataset is released as part of NVIDIA NeMo Gym, a framework… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-web_search-mcqa.nemotron-gym-knowledge-web-search-mcqa-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-knowledge-web-search-mcqa-qwen3.5-122b-131k-opencode-traces.MNLP_M3_mcqa_dataset_qasc_cotnemotron-gym-knowledge-mcqa-qwen3.5-122b-32k-tracesdrtulu_v2_nemotron_web_search_mcqaM3_MCQA_cs_science_mathstem_mcqafinance-law-mcqa
국가법령정보센터 문서를 기반으로 skt/A.X-4.0를 활용하여 생성
math_mcqa
MATH-MCQA
A multiple choice adaptation of the MATH dataset containing 12,498 competition-level mathematics problems.
Key Statistics
Metric
Value
Total Examples
12,498
Train Split
7,498
Test Split
5,000
Categories
7 (Algebra, Intermediate Algebra, Prealgebra, Geometry, Number Theory, Counting & Probability, Precalculus)
Difficulty Levels
5 (Level 1 = Easiest, Level 5 = Hardest/Competition-level)
Format
4-option multiple choice (1 correct answer + 3… See the full description on the dataset page: https://huggingface.co/datasets/stellaathena/math_mcqa.STEM-MCQA-Synthetic-55Kmcqa-multiple-answers
Dataset Summary
EVE-mcqa-multiple-answers is a Multiple-Choice Question Answering (MCQA) dataset designed to evaluate the performance of language models in the domain of Earth Observation (EO). The dataset consists of questions related to EO concepts, technologies, and applications, each accompanied by multiple answer choices, with one or more correct answer.
Dataset Structure
Each example in the dataset contains an arbitrary number of possible choices and one or more… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/mcqa-multiple-answers.mcqa-single-answer
Dataset Summary
EVE-mcqa-single-answer is a Multiple-Choice Question Answering (MCQA) dataset designed to evaluate the performance of language models in the domain of Earth Observation (EO). The dataset consists of questions related to EO concepts, technologies, and applications, each accompanied by multiple answer choices with exactly one correct answer.
Unlike multi-answer MCQA datasets, each question in this dataset has only a single correct choice, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/eve-esa/mcqa-single-answer.
