datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Warwick-STEM
Warwick STEM Dataset (WebDataset)
A collection of 19,769 experimental scanning transmission electron microscopy (STEM) images from the University of Warwick, spanning hundreds of diverse materials projects collected between 2010 and 2018.
Dataset Description
This dataset contains experimental STEM images originally published as part of the Warwick Electron Microscopy Datasets by Jeffrey Ede. The images cover a wide range of materials and imaging conditions, making them… See the full description on the dataset page: https://huggingface.co/datasets/Stemson-AI/Warwick-STEM.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.STEM2Mat
AutoMat Benchmark: STEM Image to Crystal Structure
The AutoMat Benchmark is a multimodal dataset designed to evaluate deep‑learning systems for iDPC-STEM‑based crystal‑structure reconstruction and property prediction.
Code: https://github.com/yyt-2378/AutoMat
📁 Dataset Structure
The dataset is organized into three tiers of increasing difficulty:
benchmark/
├── tier1/
│ ├── img/ # STEM images (e.g., PNG, TIFF)
│ ├── label/ # Atomic position labels… See the full description on the dataset page: https://huggingface.co/datasets/yaotianvector/STEM2Mat.stem-diagrams
STEM Diagrams
30,325 technical diagrams (block diagrams, schematics, flowcharts, architectures)
extracted from arXiv papers across six engineering fields, each with a source
attribution and a quality score. Built by an LLM-curated pipeline and used to show
that a small frozen-feature classifier can replace the paid LLM labeling gate.
Paper: Distilling an LLM Diagram-Curation Pipeline into Local Classifiers (Adnan Abbasi, Thothica, 2026)
Code:… See the full description on the dataset page: https://huggingface.co/datasets/aeyxen/stem-diagrams.Multimodal-STEM-HLE-plus-plus
multimodal-STEM-HLE++
A high-value multimodal STEM dataset designed and empirically proven to push state-of-the-art LLMs beyond their current limits.
Explore the full multimodal-STEM-HLE++ dataset: https://go.turing.com/mm-stem-hle
Why This Dataset
Post-training with RL is now the primary driver of frontier model improvement. The bottleneck is finding data at the right difficulty for current SOTA models. MMLU is saturated (>90%). HLE, once considered unsolvable, is now… See the full description on the dataset page: https://huggingface.co/datasets/TuringEnterprises/Multimodal-STEM-HLE-plus-plus.Stembind
AVR-Bench Core
This release keeps only the core fields needed for use on Hugging Face:
image, F, R, P, s1, s2, s3, and s4.
s1-s4 are annotations for the F task.
COMPOSITE-STEMThis dataset contains the full task bundles for COMPOSITE-STEM.
tasks.json contains the task instructions while answers.json contains the corresponding answers to these tasks.
refs folder contains the reference files used for each task (if applicable)
https://arxiv.org/abs/2604.09836
multimodal-lucas
Dataset card for Multi-modal LUCAS
Dataset summary
Multi-modal LUCAS aims at being a curated vision-language dataset from LUCAS survey data and in-situ field photos. LUCAS (Land Use/Cover Area Frame statistical Survey) is a land-monitoring exercise conducted by EUROSTAT in close cooperation with the Directorate-General responsible for Agriculture, with technical support from the Joint Research Centre (JRC). The survey has been repeated every three years since 2006… See the full description on the dataset page: https://huggingface.co/datasets/stemauro/multimodal-lucas.STEM-en-ms
A Bilingual Dataset for Evaluating Reasoning Skills in STEM Subjects
This dataset provides a comprehensive evaluation set for tasks assessing reasoning skills in Science, Technology, Engineering, and Mathematics (STEM) subjects. It features questions in both English and Malay, catering to a diverse audience.
Key Features
Bilingual: Questions are available in English and Malay, promoting accessibility for multilingual learners.
Visually Rich: Questions are accompanied by figures to… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/STEM-en-ms.STEM_train_cot
STEM Image Chain-of-Thought Edit Analysis Dataset
This dataset contains AI-generated Chain-of-Thought (CoT) reasoning for STEM image editing tasks, providing step-by-step analysis of edit operations.
Dataset Structure
The dataset is organized in batches:
Total batches: 26
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Fields
Each item contains:
Source image and caption (from previous stage)
Edit command (from original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train_cot.The_dataset_of_segmented_fruit_stemsmango_fruit_stem_detection
Mango Fruit Stem Detection
This dataset provides real-world RGB images of mango fruit and stems captured in agricultural field environments across Hainan Province, China. Images were collected using handheld smartphones during the December 2024 harvest period, reflecting diverse natural conditions for object detection in orchard settings. The dataset contains 1,782 images with 13,785 bounding box annotations across 2 categories.
This dataset is indexed on… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/mango_fruit_stem_detection.STEM_train_filteredsecond_model_birds_nonbirdsSTEMBench_GPT_CoT_scoreCity_mapThis dataset contains over 600 maps images of 45 various city around the world.
For trouble-shooting with the dataset, you may use this script to identify potentially corrupted files.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
STEMEditBench_v2_mediumeurosat-demo
Dataset Card for "eurosat-demo"
More Information needed
STEM_code_with_instructions_v3STEMEditBench_v2_CoTfl-invasive-birdsStemGenBench_GPTSTEM_train
STEM Image Captions Dataset
This dataset contains AI-generated captions for STEM images.
Dataset Structure
The dataset is organized in batches:
Total batches: 26
Each batch is stored in a separate directory (batch_0000, batch_0001, etc.)
Statistics
Total items: 1255641
Successfully captioned: 1255640
Errors: 1
Skipped: 0
Loading the Dataset
You can load individual batches:
from datasets import load_dataset
# Load a specific batch
batch_0 =… See the full description on the dataset page: https://huggingface.co/datasets/JackyZhuo/STEM_train.STEMGenBench_images_cot_3000STEM_code_with_instructionsSTEMBench_cotSTEMGenBench_images_3000StemBench_CoT_scoredenoise_judgingStemEditBench
