datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.dclm-stem-filteredSTEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.Logics-STEM-SFT-Dataset-Open-5.3MElectrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
stem_zh_instruction
stem_zh_instruction
内容:STEM相关指令(gpt-3.5爬取),包含物理、化学、医学、生物学、地球科学;共计256K条。
Content: STEM related instructions (gpt-3.5 crawled), including physics, chemistry, medicine, biology, and earch science. 256K instruction data in total.
学科 / Subject
文件名 / File Name
数量 / Num
物理 / Physics
phy_50380.json
50,380
化学 / Chemistry
chem_50839.json
50,839
医学 / Medicine
med_54617.json
54,617
生物学 / Biology
bio_50282.json
50,282
地球科学 / Earth Science
earth_50068.json
50,068
总计
256,186… See the full description on the dataset page: https://huggingface.co/datasets/hfl/stem_zh_instruction.salabs-stem-deep-reasoning-cot-v13
🧪 SALabs Multi-Domain STEM Deep Reasoning & Chain-of-Thought (CoT) Corpus (v13.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,500 USD) & Instant 31.7MB Master Archive DownloadInstant download of the full lossless master package containing all 1,816 JSONL reasoning records + 13 complete uncompressed text corpora (31.72 MB uncompressed total) + commercial license certificate.
🌟 Executive Summary
The SALabs STEM Deep Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-stem-deep-reasoning-cot-v13.stem-scientific-code-sample
AxiomSet Labs STEM Scientific-Code Sample
A 30-task sample of STEM reasoning and scientific-code problems across five domains.
Domains
Biology: 6 tasks
Chemistry: 6 tasks
Materials Science: 6 tasks
Mathematics: 6 tasks
Physics: 6 tasks
Each task contains two subproblems and one main problem, including prompts, scientific background, testing templates, and reference solutions.
Files
data/sample.jsonl — one task per line; used by the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/AxiomSetLabs/stem-scientific-code-sample.MMMLU-STEM-Ko
Details
This is a subset of [openai/MMMLU].
Only the subjects related to STEM were extracted from Korean subset of MMMLU.
The included subjects are
'abstract_algebra',
'anatomy',
'astronomy',
'college_biology',
'college_chemistry',
'college_computer_science',
'college_mathematics',
'college_physics',
'computer_security',
'conceptual_physics',
'electrical_engineering',
'elementary_mathematics',
'high_school_biology',
'high_school_chemistry'… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/MMMLU-STEM-Ko.stem-fixJosephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details
Dataset Card for Evaluation run of Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
Dataset automatically created during the evaluation run of model Josephgflowers/Tinyllama-STEM-Cinder-Agent-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Josephgflowers__Tinyllama-STEM-Cinder-Agent-v1-details.CC-BY-STEMM-Podcast-TranscriptsMuta-STEM-100
Muta STEM 100
Private evaluation dataset containing the fixed 100-prompt Muta STEM battery:
50 mathematics prompts (M01–M50)
50 science prompts (S01–S50)
50 multiple-choice and 50 structured-written prompts
Each row contains id, title, subject, format, text, expected, source, and suite.
This is an evaluation artifact, not a training split. Publishing or training on it would contaminate future benchmark results. Model responses are not included.
Source artifact SHA256:… See the full description on the dataset page: https://huggingface.co/datasets/timiiowolabi/Muta-STEM-100.swti-stem-20kultrafine-stem-part-1swahili-text-corpus
Dataset for Swahili Text Corpus for TTS training
Overview
This dataset contains a synthetic Swahili text corpus designed for training Text-to-Speech (TTS) models. The dataset includes a variety of Swahili phonemes to ensure phonetic diversity and high-quality TTS training.
Statistics
Format: JSONL (JSON Lines)
Data Creation
The dataset was generated using OpenAI's gpt-3.5-turbo model. The model was prompted to produce Swahili sentences that are… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-text-corpus.STEMmix
STEMmix
Generated with LMDataTools using DataMix.
Samples and combines datasets from Hugging Face.
Here's a thinking process:
Analyze User Input:
Dataset Name: STEMmix
Generated by: DataMix
Sample Entries: Two examples showing conversations between "human" and "gpt". Topics include traffic flow dynamics (chaotic dynamics, car-following models, Optimal Velocity Model) and formal logic/predicate calculus (existential/universal quantifiers, biconditional introduction).… See the full description on the dataset page: https://huggingface.co/datasets/theprint/STEMmix.STEMBench-small
STEMBench-small
Sixty research tasks. Six domains. Ten problems in each.
STEMBench-small is a public exposition of scientific research challenges: precise questions, explicit research targets, and the evidence needed to make progress. It spans proofs, computation, measurement, inference, experiments and engineering validation.
Explore the problems · Selection rationale · Review rubric · Hugging Face dataset · GitHub repository
Domain
Problems
Research covered
Biology… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/STEMBench-small.palladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.stems-predict-datastemming-instructionsstemmingCC-BY-STEMM-Podcast-Transcripts-2048STEMScoredTopics-v1.0
Dataset Summary
A synthetic dataset of 5,584 topics, each rated on a 1-5 scale for its relevance to Science, Technology, Engineering, and Mathematics (STEM).
Data Fields
topic: A string representing a topic of study or research.
stemScore: A string from "1" (least STEM) to "5" (most STEM).
Potential Uses
This dataset is useful for a variety of NLP tasks:
Classification: Train a model to classify how STEM-related a given text is.
Instruction… See the full description on the dataset page: https://huggingface.co/datasets/MultivexAI/STEMScoredTopics-v1.0.STEM-AI-mtl_Electrical-engineering-vieMMLU_STEMChosen tasks:
"college_biology",
"college_chemistry",
"college_physics",
"college_computer_science",
"college_mathematics",
"computer_security",
"conceptual_physics",
"electrical_engineering",
"elementary_mathematics",
"high_school_biology",
"high_school_chemistry",
"high_school_computer_science",
"high_school_mathematics",
"high_school_physics",
"high_school_statistics",
"machine_learning",
"astronomy",
"anatomy",
"conceptual_physics",
"abstract_algebra"
textbooks_lectures_glossaries_stemwikiReDiX-Benchmark-Stem-ita
STEM Low-Overlap Retrieval Benchmark
A STEM-focused dataset for evaluating retrieval, reranking, and RAG systems under low lexical overlap and high semantic complexity.
⚠️ Design Principle — Low OverlapThis dataset intentionally reduces lexical similarity between queries and relevant chunks.High scores from keyword-based models (e.g., BM25) may indicate shortcut exploitation rather than real understanding.
Regolo.ai 🧠
This dataset's queries were generated using… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/ReDiX-Benchmark-Stem-ita.able_stem
