datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SFT-Science-v2
Dataset Description:
Nemotron-Science-v2 is a science reasoning dataset with synthetic (synthetic MCQ, RQA) and non-synthetic vendor problems and LLM-generated solutions. It comprises three domains (Physics, Biology, and Chemistry), two question formats (multiple-choice questions [MCQ] and open questions [OpenQ]), and three generation setups: chain-of-thought (CoT) reasoning without tools, Python tool usage, and search tools usage with the Tavily API.
The solutions were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Science-v2.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/TalentZHOU/hle_material_science.Nemotron-RL-Science-v1
Dataset Description:
Nemotron-RL-Science-v1 is a reinforcement learning (RL) dataset for science reasoning. Each example provides a problem, a reference answer, and a verifiable RL environment configuration (the agent prompt, the agent/verifier reference, and the answer-extraction template) so that a policy model can be trained with verifiable rewards. It covers three domains (Physics, Biology, and Chemistry), the open-question (OpenQ) format, and two generation setups:… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Science-v1.natural-science-reasoning
Natural Sciences Reasoning: the "smolest" reasoning dataset
A smol-scale open dataset for reasoning tasks using Hugging Face Inference Endpoints. While intentionally limited in scale, this resource prioritizes:
Reproducible pipeline for reasoning tasks using a variety of models (Deepseek V3, Deepsek-R1, Llama70B-Instruct, etc.)
Knowledge sharing for domains other than Math and Code reasoning
In this repo, you can find:
The prompts and the pipeline (see the config file).
The… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/natural-science-reasoning.openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-32B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-32B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-32b-n8-flattened-logprobs-k16.openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-30B-A3B-Thinking-2507 (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-30B-A3B-Thinking-2507
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-30b-a3B-thinking-2507-n8-flattened-logprobs-k16.10001-Science-Facts
10,001 Science Facts
10,000+ obscure, surprising, and verifiable science facts
The kind that make you go "wait, really?"
🔗 GitHub Repository •
📁 Download by Category
🤔 What is this?
A curated dataset of 10,003 science facts across 32 categories — from quantum physics to parasites to the history of food.
Every fact is:
Sourced — from Wikipedia, Wikidata, academic sources
Verifiable — no LLM hallucinations
Surprising — passes the "dinner party test"… See the full description on the dataset page: https://huggingface.co/datasets/Royal-lobster/10001-Science-Facts.islamic-sciences
islamlab — The Islamic Sciences Corpus
The Islamic sciences other than Qur'an and hadith, as their authors wrote
them: 4,022 works by scholars who died between the
0st and the 14th Hijri century, cut along their own chapter
and biographical-entry boundaries into 1,864,389 units
(3.41 billion characters of Arabic), each carrying the volume and
page it sits on so a quotation can be cited rather than merely produced.
Scope is Ahl al-Sunnah wa'l-Jamāʿah, and the gate is the author… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/islamic-sciences.Global-Ocean-Science-Corpus
🌊 Global-Ocean-Science-Corpus (v2.0 Curated & Cleaned)
A Highly Curated, Large-Scale Pre-Training & RAG Corpus for Deep Ocean Sciences, Marine Biology, and Oceanography
Language Note: This dataset is a 100% English-language scientific corpus (language: "en") aggregating peer-reviewed literature, deep-sea exploration dossiers, and technical oceanographic reports from leading global marine institutes.
Global-Ocean-Science-Corpus, derin okyanus bilimleri, deniz biyolojisi… See the full description on the dataset page: https://huggingface.co/datasets/tilikumotp/Global-Ocean-Science-Corpus.llama-nemotron-science-reasoning-on-canonical-think-full
Llama-Nemotron science reasoning — Delphi canonical-think (COMPLETE, no length filter)
The complete reasoning:on science split of
nvidia/Llama-Nemotron-Post-Training-Dataset, converted once into the canonical
Delphi chat-template thinking format. 708,920 rows.
Unlike the cold-start warmup slice
open-athena/llama-nemotron-science-reasoning-on-le3000tok-100k
(and its -canonical-think variant), this build applies no length cap and no subsample — every
long-CoT science example is… See the full description on the dataset page: https://huggingface.co/datasets/laion/llama-nemotron-science-reasoning-on-canonical-think-full.data-science-code-training-pool
Data science code training pool
Public questions about writing Python with numpy, pandas, matplotlib, scikit-learn, scipy, pytorch
and tensorflow, each with the code that answers it, gathered from the datasets named below at the
pinned revisions and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 339575 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/data-science-code-training-pool.Science-QnA
Science-QnA
The Science-QnA is a large-scale, high-quality science-focused dataset (~5.63M rows) curated using synthetic data generation through distillation techniques and select open-source resources. Designed to train and evaluate reasoning-capable models in science domains with emphasis on conceptual understanding, numerical problem-solving, and exam-style Q&A patterns across Physics, Chemistry, Biology, and Mathematics.
Summary
• Domain: Science, Physics… See the full description on the dataset page: https://huggingface.co/datasets/169Pi/Science-QnA.openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16
OpenThoughts-4 Science SDG: Qwen3-4B (n=8, top-16 logprobs)
Synthetic generations from
Qwen/Qwen3-4B
on the Marin OpenThoughts-4 science SDG prompt
set.
Each prompt is sampled n=8 times, and for every generated token the dataset
stores the chosen-token log probability plus the top-16 log probabilities
over the vocabulary, enabling distillation, KL-style fine-tuning,
reranking, and uncertainty analysis.
Generation setup
Field
Value
Generator model… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/openthoughts4-science-26041-prompts-qwen3-4b-n8-flattened-logprobs-k16.science-tool-use-conversations
Science Tool-Use Conversations
This dataset contains 11,405 synthetic conversations about science questions. GLM-5.3 generated both the user and assistant messages. The assistant could run commands in shellsim, an in-memory shell and Python simulator. Each row includes a system message, the user-visible conversation, a tool-call transcript, and the tool definition. Some conversations contain no tool calls.
The questions come from the so_openq split of… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/science-tool-use-conversations.IEEE2026_BigData_MAS-4-Science-Matching
SciAgentTrace
An execution-layer trace resource for scientific-agent workload characterization.
A protocol fixes who reasons, what each role can see, when feedback returns, and
when a workflow stops. Those choices determine the sequence of model requests
that produces an answer, so protocol design is also workload design. Two
workflows that consume similar token totals can issue very different request
sequences. SciAgentTrace records that difference.
The matched core runs the… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/IEEE2026_BigData_MAS-4-Science-Matching.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.GLM-5.2-Science
GLM-5.2 · Science-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Physics · Chemistry · Biology
Token Count: 160M
Theres prompt overlap with my Kimi K2.5 dataset science subset, which I think those prompts are getting used in alot of places now
You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own.
hi - ianncity
data-science-en-id
Data Science EN-ID Parallel Corpus (Scientific Domain)
Dataset Description
This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language.
Primary Languages: English (EN) and Indonesian (ID)
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.tubitak-science-olympiad-tr
TUBITAK Science Olympiad Dataset
This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/tubitak-science-olympiad-tr.snowball-step38-science-rlvr-sft-2026-09
Snowball Step38 science RLVR SFT
The 2026.09.21-v1 directory contains the three physical 100B-token packed science mixes prepared for the balanced, proof-first, and science-forward SFT arms. The evaluated checkpoints were reached after roughly 27B scheduled token positions per arm. The glm53-rlvr1-32k-2026.09.25-v1 directory is the 33,109-conversation RLVR1 chat add-on packed into 7,036 sequences. 2026.09.25-v1 contains the three immutable manifests that combine each science mix… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/snowball-step38-science-rlvr-sft-2026-09.hle_material_science
HLE Material Science: A Specialized Benchmark for Materials Science
A Materials Science Subset of Humanity's Last Exam (HLE)
Overview
HLE Material Science is a carefully curated materials science subset derived from the Humanity's Last Exam (HLE) dataset, containing 106 high-quality expert-level questions covering 25+ materials science subfields, with 97% of questions rated as high confidence.
This dataset is designed to evaluate large language models'… See the full description on the dataset page: https://huggingface.co/datasets/stonelight/hle_material_science.task047_miscellaneous_answering_science_questions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task047_miscellaneous_answering_science_questions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task047_miscellaneous_answering_science_questions.Cannabis_Science_Data
Cannabis Science Literature QA Dataset
This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge.
Dataset Details
Dataset Description
This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/KellanF89/Cannabis_Science_Data.Dense-Information-Science-Physics-Dataset
Dense Information With Multiple Fine-tuned Variations
This dataaset has multiple for each input to learn how to express the same answer in different ways
Dataset Structure
The dataset contains two columns:
Column
Description
input
A science or quantum-physics question
output
A conversational answer to the question
Example:
{
"input": "What is quantum entanglement?",
"output": "Quantum entanglement is when two quantum systems share one… See the full description on the dataset page: https://huggingface.co/datasets/StarpowerTechnology/Dense-Information-Science-Physics-Dataset.Mixture-Science
RecursiveMAS Mixture-Science
Project Page | Code | Paper
We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Mixture-Style setting.
Dataset Details
Item
Description
Dataset
RecursiveMAS/Mixture-Science
Original file
Mixture-Science.json
Collaboration style
Mixture-Style
Used for
science specialist inner agent training
Split
train
Rows… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Mixture-Science.k-10-science-2-60k
k-10-science-2-60k
Synthetic, original expository text aligned to the K-10 (CBSE/NCERT-style) curriculum.
Rows: 60,000
Total words: 47,082,892
Subject(s): Science
Rows per grade: 1: 3,744, 10: 8,336, 2: 3,744, 3: 6,876, 4: 5,481, 5: 5,131, 6: 6,384, 7: 6,528, 8: 6,824, 9: 6,952
Fields
Field
Type
Description
text
string
The generated passage
subject
string
Subject name
grade
int
Grade level
word_count
int
Number of words in text
adaption-science-seed
Science Q&A
Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers.
Rows
12,000
Domain
science
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-seed.adaption-science-biochem-seed
Science Q&A
Science questions (chemistry, physics, biology; the biochem variant focuses on life sciences) with worked answers.
Rows
12,458
Domain
science
Format
data.parquet, one row per example
Licence
apache-2.0
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-science-biochem-seed.task701_mmmlu_answer_generation_high_school_computer_science
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task701_mmmlu_answer_generation_high_school_computer_science
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task701_mmmlu_answer_generation_high_school_computer_science.NCERT_Political_Science_12th
