mcqa
mistral-7b-qlora-ai2_arc-full-mcqa-mergedmistral-7b-qlora-ai2_arc-full-mcqa-setllm-mergedmistral-7b-qlora-ai2_arc-full-mcqa-choices-mergedattention-avengers_-_Qwen1.5-0.5B-Chat-SFT-MCQA-ggufLLM-Elicitation_-_mistral-instruct-mcqa-circuit-broken-ggufLLM-Elicitation_-_mistral-mcqa-circuit-broken-ggufdrudilorenzo-mcqa_sft-GGUFOrelian-mcqa_quantized-GGUF
Datasets
All datasets matching “mcqa”copycolors_mcqaThis dataset consists of formatted n-way multiple choice questions, where n is in [2,10]. The task itself is simply to copy the prototypical color from the context and produce the corresponding color's answer choice letter.
The "prototypical colors" dataset instances themselves come from Memory Colors (Norland et al. 2021) and corypaik/coda (instances whose object_group is 0, indicating participants agreed on a prototypical color of that object).
Nemotron-RL-knowledge-mcqa
Dataset Description:
The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.med_mcqaFrom "MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering"
(Pal et al.), MedMCQA is a "multiple-choice question answering (MCQA) dataset designed to address
real-world medical entrance exam questions." The dataset "...has more than 194k high-quality AIIMS & NEET PG
entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average
token length of 12.77 and high topical diversity."
The following is an example from… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/med_mcqa.driving_mcqa
DrivingExamMCQA
The DrivingExamMCQA dataset is a Multiple-Choice Question Answering (MCQA) collection based on real driving exam questions. It supports multilingual assessment across three languages: Arabic (ar), French (fr), and English (en) (with translations).
Overview
Each language includes two modalities:
Image-supported questions (_img splits):
Questions paired with an image (e.g., road signs, traffic scenarios).
Text-only questions (_text splits):
Standard… See the full description on the dataset page: https://huggingface.co/datasets/ESmike/driving_mcqa.science-mcqa-training-pool
Science multiple-choice training pool
Public multiple-choice science questions from three datasets, read at the pinned revisions named
below and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 182035 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
question
the question text, as its source publishes it
options
the answer options, as… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/science-mcqa-training-pool.NLU-Belebele-MCQA
