datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MNLP_M2_mcqa_datasetThis dataset contains the MCQA and instruction finetuning datasets:
The messages column is used by the instruction finetuning dataset
The choices, question, context, and answer columns are used by the MCQA dataset
For the MCQA dataset (of only single answer) contains a mixture of the train, validation and test splits from this datasets as to have for training and testing:
mmlu auxiliary train we only use the stem subsets
mmlu we only use the stem subsets
ai2_arc
ScienceQA
math_qa… See the full description on the dataset page: https://huggingface.co/datasets/andresnowak/MNLP_M2_mcqa_dataset.MNLP_M2_mcqa_dataset
Dataset Card for SCP-116K
Recent Updates
We have made significant updates to the dataset, which are summarized below:
Expansion with Mathematics Data:Added over 150,000 new math-related problem-solution pairs, bringing the total number of examples to 274,166. Despite this substantial expansion, we have retained the original dataset name (SCP-116K) to maintain continuity and avoid disruption for users who have already integrated the dataset into their workflows.
Updated… See the full description on the dataset page: https://huggingface.co/datasets/asazheng/MNLP_M2_mcqa_dataset.nemotron-nano-rl-mcqa-19k
Nemotron Nano RL MCQA 19K
Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format.
Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.Synthetic_Dataset_For_MCQA
