Team Ai
Datasetpublic

Emulated-Inc/science-mcqa-training-pool

Science multiple-choice training pool Public multiple-choice science questions from three datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 182035 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file question the question text, as its source publishes it options the answer options, as… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/science-mcqa-training-pool.

sourceHugging Facecc-by-sa-4.0updated 17d agoView on Hugging Face
0likes704downloads
Dataset Card

Science multiple-choice training pool

Public multiple-choice science questions from three datasets, read at the pinned revisions named below and laid out twice. Train on either layer or on both.

pool.jsonl

Every source rewritten into one shape, 182035 rows, one JSON object per line, with these fields.

FieldWhat it holds
ida row identifier unique within this file
questionthe question text, as its source publishes it
optionsthe answer options, as a list of strings
answerthe letter of the correct option, A for the first option listed
subjectthe subject label the source gives the row, or unknown
languagethe language of the row
sourcethe name of the source directory the row came from
source_repo, source_revisionthe dataset and the revision it was read at
source_subset, source_split, source_idwhere the row sits in that set
provenance_classhow the row came to exist
licencethe licence of the source it came from

The options of every row were permuted once when this file was built, under the recorded seed 20260911, and the answer letter is written in the permuted order, so that a source which publishes its correct option in a fixed position cannot teach that position. Rows are deduplicated across sources on the question together with its option set, keeping the first source that carries the question in the order of the sections below.

sources/

The same data untouched, 200550 rows, one directory per source, holding the files at the paths, in the parquet format and with the columns its own repository publishes. Nothing here was renamed, reshaped, reordered or deduplicated. Use this layer if you want a field the rewritten one drops, such as the written explanations and topic names in MedMCQA, or if you would rather order the options yourself.

The sources

sources/medmcqa

Indian medical entrance examination questions, AIIMS and NEET PG, 1991 to the present, with the examining body's own answer key. From openlifescienceai/medmcqa at revision 91c6572c454088bf71b679ad90aa8dffcd0d5868, files data/train-00000-of-00001.parquet, data/validation-00000-of-00001.parquet. 185732 rows here, of which 167356 also appear in pool.jsonl. Provenance class human, licence apache-2.0. Its own fields are id, question, opa to opd for the four options, cop for the zero-based index of the correct one, choicetype, exp for a written explanation, subjectname, topic_name.

Worth knowing. The apache-2.0 tag covers the packaging rather than the examination boards' copyright in the underlying questions, which the dataset card does not address.

sources/ai2_arc

Grade-school science examination questions written for human tests, in an easy and a challenge partition. From allenai/ai2_arc at revision 210d026faf9955653af8916fad021475a3f00453, files ARC-Easy/train-00000-of-00001.parquet, ARC-Easy/validation-00000-of-00001.parquet, ARC- Challenge/train-00000-of-00001.parquet, ARC-Challenge/validation-00000-of-00001.parquet. 4185 rows here, of which 4184 also appear in pool.jsonl. Provenance class human, licence cc-by- sa-4.0. Its own fields are id, question, choices with parallel text and label lists, answerKey.

Worth knowing. These items are widely present in web pretraining corpora. Infini-gram mini reports a dirty rate of 34.10 percent for the challenge partition and 31.70 percent for the easy one against DCLM-baseline, so a model pretrained on web text may have met them already.

sources/exams

High-school examination questions collected from the official state examinations of several ministries of education, over 24 subjects. From mhardalov/exams at revision 4ff10804abb3341f8815cacd778181177bba7edd, files multilingual/train-00000-of-00001.parquet, multilingual/validation-00000-of-00001.parquet. 10633 rows here, of which 10495 also appear in pool.jsonl. Provenance class human, licence cc-by-sa-4.0. Its own fields are id, question with a stem and choices holding parallel text, label and para lists, answerKey, and info with the grade, the subject and the language.

Worth knowing. The multilingual configuration only, and its own test split is left out. The subjects run past the sciences into history, philosophy and business, and the subject field is how to select. No row is in English.

Provenance and licences

Every row was written by a person. All three sources are questions written for human examinations, which is the provenance class human. Nothing in the pool was generated by a model.

The pool as a whole is offered under cc-by-sa-4.0, which is the most restrictive term its sources compose to. The sources themselves are apache-2.0 for MedMCQA and cc-by-sa-4.0 for ARC and for EXAMS, and every one of them permits commercial use. Each rewritten row carries its own in the licence field and each directory under sources/ is one source, so a subset under a single licence can be selected. Attribution for the share-alike sources goes to the Allen Institute for AI for ARC and to the authors of EXAMS, and questions taken from national examinations remain the property of the boards that wrote them.

Filtering

Rows whose question duplicated or closely paraphrased a question in a held-out evaluation set were removed before publication, from both layers alike, by a word 8-gram overlap check (1257 rows) followed by an embedding similarity check (70 rows). That evaluation set is not distributed here. Nothing else was filtered: no subject, no language and no difficulty was selected for or against, so the pool still holds the history, philosophy and business questions its multi-subject sources carry alongside the sciences, and the subject field is how to select what you want.