datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.STEM
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.eai-taxonomy-stem-w-dclm-100b-sample
🔬 EAI-Taxonomy STEM w/ DCLM (100B sample)
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.dclm-stem-filteredstemdata
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/si-m07/stemdata.Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.Natural-Reasoning-STEM-25Kshreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text (CPT Healing Corpus)
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentation across 2,400+ partitioned Parquet shards.
🔬 Architectural Role in Continual Pre-Training (CPT) Healing
This corpus served as the foundational Continual Pre-Training (CPT)… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.STEM_COTstem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.pat-stem-pretrainChina-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.Logics-STEM-SFT-Dataset-Open-5.3Mstem-corpusMMLU-STEMThis contains a subset of STEM subjects defined in MMLU by the original paper.
The included subjects are
'abstract_algebra',
'anatomy',
'astronomy',
'college_biology',
'college_chemistry',
'college_computer_science',
'college_mathematics',
'college_physics',
'computer_security',
'conceptual_physics',
'electrical_engineering',
'elementary_mathematics',
'high_school_biology',
'high_school_chemistry',
'high_school_computer_science',
'high_school_mathematics',
'high_school_physics'… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-STEM.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.ki-qwen-instruct-synthetic_1_stem_only-sft-temp0.6-on-mmlu_pro-0shot_cot-scillm-da553bdec9STEM_sftSTEM_DPOElectrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
stem_zh_instruction
stem_zh_instruction
内容:STEM相关指令(gpt-3.5爬取),包含物理、化学、医学、生物学、地球科学;共计256K条。
Content: STEM related instructions (gpt-3.5 crawled), including physics, chemistry, medicine, biology, and earch science. 256K instruction data in total.
学科 / Subject
文件名 / File Name
数量 / Num
物理 / Physics
phy_50380.json
50,380
化学 / Chemistry
chem_50839.json
50,839
医学 / Medicine
med_54617.json
54,617
生物学 / Biology
bio_50282.json
50,282
地球科学 / Earth Science
earth_50068.json
50,068
总计
256,186… See the full description on the dataset page: https://huggingface.co/datasets/hfl/stem_zh_instruction.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.Multimodal-STEM-HLE-plus-plus
multimodal-STEM-HLE++
A high-value multimodal STEM dataset designed and empirically proven to push state-of-the-art LLMs beyond their current limits.
Explore the full multimodal-STEM-HLE++ dataset: https://go.turing.com/mm-stem-hle
Why This Dataset
Post-training with RL is now the primary driver of frontier model improvement. The bottleneck is finding data at the right difficulty for current SOTA models. MMLU is saturated (>90%). HLE, once considered unsolvable, is now… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Multimodal-STEM-HLE-plus-plus.stem_mcqastem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.STEM-MCQA-Synthetic-55KStemQAMixture
StemQAMixture
Science QA dataset with four subjects: biology, chemistry, math, and physics.
Usage
from datasets import load_dataset
# Load a specific subject
ds = load_dataset("4gate/StemQAMixture", "math", split="train")
# Load all subjects
ds_bio = load_dataset("4gate/StemQAMixture", "biology")
ds_chem = load_dataset("4gate/StemQAMixture", "chemistry")
Schema
question: The question text
answer: The answer text
reference_answer: Optional reference… See the full description on the dataset page: https://huggingface.co/datasets/4gate/StemQAMixture.
