datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.eai-taxonomy-stem-w-dclm
🔬 EAI-Taxonomy STEM w/ DCLM
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 1742 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that require… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm.Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.melodix-stemsSTEM
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/stemdataset/STEM.MASTER-STEM-DATASETeai-taxonomy-stem-w-dclm-100b-sample
🔬 EAI-Taxonomy STEM w/ DCLM (100B sample)
🏆 Website | 🖥️ Code | 📖 Paper
A high-quality STEM dataset curated from web data using taxonomy-based filtering, containing 100 billion tokens of science, technology, engineering, and mathematics content.
🎯 Dataset Overview
This dataset is part of the Essential-Web project, which introduces a new paradigm for dataset curation using expressive metadata and simple semantic filters. Unlike traditional STEM datasets that… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/eai-taxonomy-stem-w-dclm-100b-sample.dclm-stem-filteredstemdata
STEM Dataset
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] • 📽 [Slides] • 📋 [Poster]
This dataset is proposed in the ICLR 2024 paper: Measuring Vision-Language STEM Skills of Neural Models. We introduce a new challenge to test the STEM skills of neural models. The problems in the real world often require solutions, combining knowledge from STEM (science, technology, engineering, and math). Unlike existing datasets, our dataset requires the understanding of… See the full description on the dataset page: https://huggingface.co/datasets/si-m07/stemdata.Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.Natural-Reasoning-STEM-25Kstemshreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text (CPT Healing Corpus)
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentation across 2,400+ partitioned Parquet shards.
🔬 Architectural Role in Continual Pre-Training (CPT) Healing
This corpus served as the foundational Continual Pre-Training (CPT)… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.STEM2Mat
AutoMat Benchmark: STEM Image to Crystal Structure
The AutoMat Benchmark is a multimodal dataset designed to evaluate deep‑learning systems for iDPC-STEM‑based crystal‑structure reconstruction and property prediction.
Code: https://github.com/yyt-2378/AutoMat
📁 Dataset Structure
The dataset is organized into three tiers of increasing difficulty:
benchmark/
├── tier1/
│ ├── img/ # STEM images (e.g., PNG, TIFF)
│ ├── label/ # Atomic position labels… See the full description on the dataset page: https://huggingface.co/datasets/yaotianvector/STEM2Mat.STEM_COTstem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.pat-stem-pretrainChina-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.Logics-STEM-SFT-Dataset-Open-5.3Mstem-corpusMMLU-STEMThis contains a subset of STEM subjects defined in MMLU by the original paper.
The included subjects are
'abstract_algebra',
'anatomy',
'astronomy',
'college_biology',
'college_chemistry',
'college_computer_science',
'college_mathematics',
'college_physics',
'computer_security',
'conceptual_physics',
'electrical_engineering',
'elementary_mathematics',
'high_school_biology',
'high_school_chemistry',
'high_school_computer_science',
'high_school_mathematics',
'high_school_physics'… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-STEM.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.vocalsetki-qwen-instruct-synthetic_1_stem_only-sft-temp0.6-on-mmlu_pro-0shot_cot-scillm-da553bdec9Stems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source… See the full description on the dataset page: https://huggingface.co/datasets/drizzymedia/Stems-Evaluation-Kit.STEM_sfttushe-grade-school-stem
Tushe Community Grade School STEM
Open dataset of grade-school STEM (Science, Technology, Engineering, Mathematics) textbooks, curated for Tushe Community and aligned with curriculum use (e.g. CAPS-aligned content).
Data Fields (per book JSON)
Field
Type
Description
source_file
string
Original .txt filename
title
string
Derived book title (e.g. "Grade 8A Mathematics")
table_of_contents
list
[{ "section_id", "title" }, ...]
front_matter
string
Intro… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/tushe-grade-school-stem.STEM_DPO
