datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text (CPT Healing Corpus)
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentation across 2,400+ partitioned Parquet shards.
🔬 Architectural Role in Continual Pre-Training (CPT) Healing
This corpus served as the foundational Continual Pre-Training (CPT)… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.stem-reasoning-complex
STEM-Reasoning-Complex: High-Fidelity Scientific CoT Dataset
1. Dataset Summary
STEM-Reasoning-Complex is a curated collection of 118.255 high-quality samples designed for Supervised Fine-Tuning (SFT) and alignment of Large Language Models. The dataset focuses on four core disciplines: Biology, Mathematics, Physics, and Chemistry.
Unlike standard QA datasets, each entry provides a structured Chain-of-Thought (CoT) reasoning process, enabling models to learn… See the full description on the dataset page: https://huggingface.co/datasets/galaxyMindAiLabs/stem-reasoning-complex.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740,000+… See the full description on the dataset page: https://huggingface.co/datasets/lfaviate/China-K12-STEM-10K-CoT-Reasoning.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.tushe-grade-school-stem
Tushe Community Grade School STEM
Open dataset of grade-school STEM (Science, Technology, Engineering, Mathematics) textbooks, curated for Tushe Community and aligned with curriculum use (e.g. CAPS-aligned content).
Data Fields (per book JSON)
Field
Type
Description
source_file
string
Original .txt filename
title
string
Derived book title (e.g. "Grade 8A Mathematics")
table_of_contents
list
[{ "section_id", "title" }, ...]
front_matter
string
Intro… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/tushe-grade-school-stem.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
AI-Research-Evaluation-Repository-STEM
AI-STEM-Research-Eval-Dataset
Overview
This dataset contains AI-generated scientific reports across STEM domains, accompanied by structured metadata, prompt documentation, reference validation, and hallucination annotations.
It is designed as an open research resource to study the capabilities, limitations, and reliability of large language models (LLMs) in generating scientific content.
The dataset enables systematic analysis of how AI systems perform in… See the full description on the dataset page: https://huggingface.co/datasets/sreearravind/AI-Research-Evaluation-Repository-STEM.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.stem-reasoning-ccbysa-011
YouAI Data — stem-reasoning-ccbysa-011
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 11 step-by-step reasoning chains and 989 instruction/response pairs across 454 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 563 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-011.Punjabi-STEM-Frontier-CoT-Corpus
⚠️ Provenance note (2026-10-06). Solutions were generated with AI assistance and not reviewed by subject experts; check the reasoning before using it for teaching or as ground truth.
ੴ Punjabi STEM & Frontier Chain-of-Thought (CoT) Corpus
☬ ਪੰਜਾਬੀ (ਗੁਰਮੁਖੀ) ਵਿਗਿਆਨ ਅਤੇ ਉੱਚ-ਗਣਿਤ ਕਦਮ-ਦਰ-ਕਦਮ ਤਰਕ ਡਾਟਾਸੈੱਟ
👨💻 Research & Engineering Architecture
Architect & Developer: Gurpreet Singh Dhillon (Nam-toon Studio)
Vision: Sovereign Indic… See the full description on the dataset page: https://huggingface.co/datasets/Nam-toon-studio/Punjabi-STEM-Frontier-CoT-Corpus.STEM-prompts-1M
Generated with GPT-OSS-120B on medium reasoning, 1 million stem prompts in this distribution:
50% Programming
30% Science
20% Math
stem-reasoning-v1.0.0-ccbysa-001
YouAI Data — stem-reasoning-v1.0.0-ccbysa-001
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 394 step-by-step reasoning chains and 569 instruction/response pairs across 332 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 364 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-001.stem-reasoning-ccbysa-005
YouAI Data — stem-reasoning-ccbysa-005
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 248 step-by-step reasoning chains and 729 instruction/response pairs across 696 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 559 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-005.stem-reasoning-ccbysa-002
YouAI Data — stem-reasoning-ccbysa-002
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 529 step-by-step reasoning chains and 439 instruction/response pairs across 591 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 408 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-002.stem-tr-instruct-1k
Eding STEM TR Instruct 1K
Türkçe K-12 STEM ve kodlama eğitimi için instruction-tuning veri seti.
Veri Seti Hakkında
Bu veri seti, Türkiye'deki K-12 seviyesinde STEM ve kodlama eğitimi için
hazırlanmış 1.000 instruction-output çiftinden oluşur.
Kategoriler
Arduino: LED, sensör, motor projeleri, devre tasarımı
Scratch: Blok tabanlı programlama, oyun yapımı, animasyon
mBlock: mBot robot programlama, sensör kullanımı
Robotik: PID kontrol, çizgi izleme… See the full description on the dataset page: https://huggingface.co/datasets/alimkacar/stem-tr-instruct-1k.stem-reasoning-ccbysa-10k-001
stem-reasoning-ccbysa-10k-001
Combined 10K CC-BY-SA STEM reasoning dataset — a sequential merge of ten 1K sets.
Total examples: 10000
License: CC-BY-SA-4.0
Mean quality score: 4.411 (min 4.25, max 4.81)
Unique sources: 2387
DPO pairs: 3650
Difficulty: {'advanced': 2774, 'introductory': 5409, 'intermediate': 1817}
Products: {'reasoning_chain': 1531, 'instruction_pair': 8377, 'code_instruction': 92}
Combined from
stem-reasoning-ccbysa-001… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-10k-001.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.stem-reasoning-ccbysa-006
YouAI Data — stem-reasoning-ccbysa-006
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 251 step-by-step reasoning chains and 729 instruction/response pairs across 662 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 266 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-006.shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.omni-stem
GitHub
Website
Paper (Coming Soon)
Dataset Details
This dataset is a combination of all datasets used for expert finetuning in a curriculum order of Math, Science, Technology, and Engineering based on similarities between each domain’s fundamentals, aiding model learning. This dataset contains a mixture of chat-based and continued-pretrain data.
Sources
This dataset was sourced from the following open-sourced datasets:
Math
meta-math/MetaMathQA… See the full description on the dataset page: https://huggingface.co/datasets/omniomni/omni-stem.hausa-stem-reasoning-with-cultural-context
Hausa STEM Reasoning with Cultural Context
Abstract
We present the first large-scale bilingual Hausa-English STEM reasoning dataset with deep cultural adaptation, containing 2,640 high-quality question-answer pairs translated from the STEM-Reasoning-Complex dataset. Our work introduces the "Shehin Malamin Kimiyya" (The Wise Scholar of Science) translation framework, which transforms Western scientific concepts into culturally-embedded Hausa explanations using systematic… See the full description on the dataset page: https://huggingface.co/datasets/Tushe/hausa-stem-reasoning-with-cultural-context.stem-reasoning-ccbysa-013
YouAI Data — stem-reasoning-ccbysa-013
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 999 instruction/response pairs across 394 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 549 DPO preference pairs as a free companion dataset. Includes 994 raw reasoning examples… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-013.stem-reasoning-ccbysa-007
YouAI Data — stem-reasoning-ccbysa-007
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 21 step-by-step reasoning chains and 979 instruction/response pairs across 674 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 585 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-007.stem-reasoning-ccbysa-008
YouAI Data — stem-reasoning-ccbysa-008
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 73 step-by-step reasoning chains and 921 instruction/response pairs across 667 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 361 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-008.stem-reasoning-v1.0.0-ccbysa-002
YouAI Data — stem-reasoning-v1.0.0-ccbysa-002
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 239 step-by-step reasoning chains and 723 instruction/response pairs across 443 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 476 DPO preference pairs as a free companion… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-v1.0.0-ccbysa-002.wikipedia_stem_small_rag_embeddings
STEMWikiSmallRAG with embeddings
This dataset contains wikipedia entries from STEM field, unfortunately there is also Business&Economics... but I thought it may contain some useful data as well, even by accident.
Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_stem_small_rag_embeddings.STEM_train_cot
STEM Image Chain-of-Thought Edit Analysis Dataset
This dataset contains AI-generated Chain-of-Thought (CoT) reasoning for STEM image editing tasks, providing step-by-step analysis of edit operations.
Dataset Structure
The dataset is organized in batches:
Total batches: 26
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
Source image and caption (from previous stage)
Edit command (from original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_cot.quadmix-stem-v1
QuaDMix-STEM v1: STEM-Focused Proxy Validation Set
Script: scripts/validation_set/prepare_stem_v1.py
HuggingFace: liujin99/quadmix-stem-v1
Files: stem_v1_tokenized.pt, stem_v1.parquet
Overview
STEM v1 is a validation set designed to focus the proxy model's optimization signal on STEM capabilities — mathematics, science knowledge, and logical reasoning. Unlike CAP v1 (broad capability coverage) or core_bmk (benchmark test format), STEM v1 uses only tasks that… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/quadmix-stem-v1.stem-reasoning-ccbysa-003
YouAI Data — stem-reasoning-ccbysa-003
Dataset Description
YouAI Data — 1,000 STEM training examples extracted from verified CC-BY-SA expert sources — real domain experts solving real problems, not synthetic LLM generation. Contains 1,000 instruction/response pairs across 486 unique sources. Every example traces to a source URL, available source metadata, and verified license. Includes 101 DPO preference pairs as a free companion dataset. Includes 732 raw reasoning… See the full description on the dataset page: https://huggingface.co/datasets/YouAIData/stem-reasoning-ccbysa-003.
