datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fashion_mnist
Dataset Card for FashionMNIST
Dataset Summary
Fashion-MNIST is a dataset of Zalando's article images—consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes. We intend Fashion-MNIST to serve as a direct drop-in replacement for the original MNIST dataset for benchmarking machine learning algorithms. It shares the same image size and structure of training and testing… See the full description on the dataset page: https://huggingface.co/datasets/zalando-datasets/fashion_mnist.TALKtoME
TALKtoME: Educational Materials for Speech and Language Acquisition in Autism
This dataset contains image and video samples of action verbs and verb+noun pairs. It is designed to support machine learning tasks related to visual understanding, action recognition, and language grounding.
Dataset Structure
The dataset contains the following main folders:
action_verbs/images/: image samples organized by action verb categories.
action_verbs/videos/: video samples organized by… See the full description on the dataset page: https://huggingface.co/datasets/LSL-datasets/TALKtoME.cctv-datasets
CCTV Datasets for helmet detection + ANPR
Training and evaluation data used by vivekvar/helmet-v5 and vivekvar/helmet-v4.
Source: Andhra Pradesh RTGS CCTV feeds (public road cameras). All crops and frames are from motorcycle traffic scenes.
Folders
Folder
Contents
Purpose
merged_v3/
YOLO-format dataset (data.yaml + train/valid/test)
Bike + rider detection training
clean_merged_data/
Cleaned / deduped crop set
Base training data for v4
extra_khadatkar/… See the full description on the dataset page: https://huggingface.co/datasets/vivekvar/cctv-datasets.sketch2stl-datasets
sketch2stl datasets
CMU 24-679 Project 1 · Serena Sun & PK. Source CAD data: Fusion 360 Gallery Reconstruction r1.0.1 (Autodesk, license included).
Folder
For
Manual (counts toward the ≥500)
Synthetic / derived (training only)
ML1_stroke_recognizer/
ML 1: what kind of stroke is this? (line, arc, circle, rect, polyline) · trained from scratch
325 hand-drawn + 197 human-reviewed Fusion strokes = 522
29,387 synthetic strokes
ML2_smart_suggestions/
ML 2: what should… See the full description on the dataset page: https://huggingface.co/datasets/pakiino/sketch2stl-datasets.HSI_Datasets
🛰️ HSI Datasets Collection
A comprehensive collection of 24 publicly available Hyperspectral Image (HSI) datasets curated for research in hyperspectral image classification, land use/land cover (LULC) mapping, and remote sensing deep learning benchmarks.
Total Size: ~20.1 GBLicense: Apache 2.0Maintained by: Tanishq Rachamalla, Aryan DasHuggingFace Dataset: https://huggingface.co/datasets/Tanishq165/HSI_DatasetsPaper: Hyperspectral Image Models: Technical Report… See the full description on the dataset page: https://huggingface.co/datasets/Tanishq165/HSI_Datasets.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.dermalens-datasets
DermaLens Skin Cancer Dataset
This dataset repo documents the data pipeline used to train the DermaLens V3 skin cancer classification model.
Source Dataset
HAM10000 (Human Against Machine with 10000 training images) — accessed via marmal88/skin_cancer on HuggingFace.
from datasets import load_dataset
ds = load_dataset("marmal88/skin_cancer")
Dataset Statistics
Split
Images
Malignant
Benign
Positive Rate
Train
10,683
~2,093
~8,590
19.6%… See the full description on the dataset page: https://huggingface.co/datasets/dheraingoud/dermalens-datasets.clide_synthetic_datasets
CLIDE Synthetic Image Datasets
📄 Paper • 💻 Code • 🌐 Webpage • 🎥 Video
A collection of synthetic images generated by modern text-to-image models, organized by domain and generator.
The dataset is designed to support analysis and evaluation of generated-image detection methods under domain and generator shifts.
🗂️ Dataset Structure
The dataset contains two visual domains:
💥🚗 Damaged Cars
Synthetic images of damaged cars generated by multiple… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/clide_synthetic_datasets.Stomatal_Images_DatasetsThis new dataset is designed to solve image classification and segmentation tasks and is crafted with a lot of care.tornadonet-datasets
TornadoNet Dataset: Street-View Building Damage Assessment
Dataset Description
TornadoNet is a comprehensive benchmark dataset for automated post-disaster building damage assessment using street-level imagery. The dataset contains high-resolution geotagged images collected following the December 10-11, 2021 Midwest U.S. tornado outbreak, providing realistic conditions for evaluating modern object detection architectures on multi-level damage classification tasks.… See the full description on the dataset page: https://huggingface.co/datasets/crumeike/tornadonet-datasets.ai-auto-train-datasets-cuda-5d-quantum-mindmap-simulations-generator-zkevms-immutablexPrompt_Tuning_Datasets_with_Foreground
⭐ Dataset Introduction
The standard datasets (except ImageNet) used for CLIP-based Prompt Tuning research (e.g., CoOp).
Based on the original datasets, this repository adds foreground segmentation masks (generated by SEEM) of all raw images.
For the foreground masks, the RGB value of the foreground region is [255, 255, 255], and the background region is [0, 0, 0].
The shorter side is always fixed to 512 px, and the scaling ratio is the same as that of the… See the full description on the dataset page: https://huggingface.co/datasets/JREion/Prompt_Tuning_Datasets_with_Foreground.Synset-Background-Effect-Datasets
Synset Background Effect Datasets
For investigating the effect of background on feature importance and classification performance, we systematically generated six synthetic datasets for the
task of traffic sign recognition, which differ only in their degree of camera variation and background correlation. Each of these datasets contains 82 classes
of traffic signs with 1,100 images per class, resulting in 90,200 images per dataset, summing up to a total of 541,200 images.… See the full description on the dataset page: https://huggingface.co/datasets/FraunhoferIOSB/Synset-Background-Effect-Datasets.cutclean-datasets
CutClean: balanced datasets
Data accompanying:
Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise, Enzo Tartaglione.
CutClean: Neural Network Pruning for Privacy-Preserving Inference.
Pattern Recognition — Proceedings of the 28th International Conference on Pattern
Recognition (ICPR 2026), Lyon, France. Lecture Notes in Computer Science, Springer
Nature Switzerland, pp. 450–465.
doi:10.1007/978-3-032-31452-9_30
Code: https://github.com/MaglioloLeonardo/CutClean
The… See the full description on the dataset page: https://huggingface.co/datasets/imDalton/cutclean-datasets.FSL_datasetsspatial-datasets
Please note, since the dataset is under heavy development, any access requests won't be accepted at this moment.
2D_Shape_Image_Datasetsmicroscopy-datasets-index
Microscopy Datasets Index
I kept losing track of which microscopy datasets exist and what format they're in. So I made an index. 300+ open datasets, searchable by domain, task, microscopy type, and license.
What's in here
A single reference file (JSONL and Parquet) cataloging 309 publicly available microscopy datasets. Each entry includes:
id — short slug
name — human-readable dataset name
source — who published it (Broad Institute, Kaggle, ISBI, etc.)
url — direct link… See the full description on the dataset page: https://huggingface.co/datasets/Laborator/microscopy-datasets-index.kaeva-deepfake-datasets
Kaeva Deepfake Detection — Training Datasets (V1–V9)
This repository documents all training datasets used across Kaeva deepfake detection model versions V1 through V9. No raw data is hosted here — this serves as a comprehensive reference card.
Training code: Viraj-FG/kaeva-verify/training/
Dataset Inventory
Established Benchmarks
Dataset
Type
Source
License
CIFAKE
Real + AI-generated (CIFAR-10 scale)
HF: Bird/CIFAKE
CC BY-SA 4.0
ArtiFact… See the full description on the dataset page: https://huggingface.co/datasets/Vi0509/kaeva-deepfake-datasets.datasetsss
