datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MNIST_train
IllusionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. The dataset is intended for training models to recognize MNIST digits embedded as visual illusions (pareidolia) in generated scenes and to reject images that contain no illusion.
MNIST source-condition images were sampled and resized to 512 × 512 pixels, combined… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_train.Language-Grounded_Sparse_Encoder_Training
Language-Grounded Sparse Encoder (LanSE) — Training Data
This repository hosts the AI-generated images and human annotation datasets accompanying the paper:
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
Yiming Tang, Arash Lagzian, Srinivas Anumasa, Qiran Zou, Yingtao Zhu, Ye Zhang, Trang Nguyen, Yih-Chung Tham, Ehsan Adeli, Ching-Yu Cheng, Yilun Du, Dianbo Liu
National University of Singapore · Tsinghua University · Stanford University ·… See the full description on the dataset page: https://huggingface.co/datasets/DesmondYMTang2024/Language-Grounded_Sparse_Encoder_Training.FashionMnist_train
IllusionFashionMNIST — Training Set
Dataset summary
This repository contains the training split of IllusionFashionMNIST, one of the four datasets introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It is designed to train and evaluate models on the recognition of Fashion-MNIST categories embedded as visual illusions (pareidolia) in generated scenes.
The source-condition images are sampled from Fashion-MNIST and resized to… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_train.IllusionAnimals_train
IllusionAnimals — Training Set
Dataset summary
This repository contains the training split of IllusionAnimals, one of the four benchmarks introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. It supports training models to identify animal categories embedded as visual illusions (pareidolia) in generated scenes and to recognize when no illusion is present.
The source-condition animal images were generated with SDXL-Lightning.… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_train.OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World.
Code: https://github.com/iamwangyabin/OpenSDI
gui-odyssey-train
Dataset Card for GUI Odyssey (Train Split)
⬆️ Test split shown above, but this also represents the train split.
This is a FiftyOne dataset with 89365 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/gui-odyssey-train")
# Launch… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/gui-odyssey-train.dronescapes2_annotated_train_set
Dataset Card for DroneScapes2 (annotated train set)
This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.AVM_Segmentation_train
Dataset Card for AVM (Around View Monitoring) Semantic Segmentation Dataset
This repository provides a FiftyOne-compatible version of the AVM semantic segmentation dataset for autonomous parking systems, with enhanced metadata and visualization capabilities.
This is a FiftyOne dataset with 6763 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/AVM_Segmentation_train.NTIRE-RobustAIGenDetection-train
Training set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
📄 CVPR 2026 Workshop Paper · NTIRE 2026
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-train.open-vision-banana-snvc-train-full
SNVC-50M v5_full — Multi-Task Vision Dataset
Description
This dataset is a curated subset of the SenseNova Vision Corpus 50M (SNVC-50M), containing 43,509 samples across 6 vision task families and 31 source datasets. Each sample follows a conversational format with interleaved <image> tokens, designed for training vision-language models (VLMs).
Coverage: 43,509 / 57,878 (75.2%) of the original sampling plan. 23 datasets at 100%, 8 partial, 12 unrecoverable… See the full description on the dataset page: https://huggingface.co/datasets/gatilin/open-vision-banana-snvc-train-full.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.BioTrove-Train
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
Description
See the BioTrove dataset card on HuggingFace to access the main BioTrove dataset (161.9M)
BioTrove comprises well-processed metadata with full taxa information and URLs pointing to image files. The metadata can be used to filter specific categories, visualize data distribution, and manage imbalance effectively. We provide a collection of software… See the full description on the dataset page: https://huggingface.co/datasets/BGLab/BioTrove-Train.trash_bin_train
Trash Bin Train Dataset
This dataset contains real-world images of trash and waste bins, compiled from the Roboflow Universe dataset trash-train-vz2jm.
Details
Total Images: 1059
Format: Flat folder containing trash_bin_train_xxxxx.jpg
Use Case: Training object detection and classification models to identify trash bins and waste management issues.
ur5fail_train_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.bdv2fail_train_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.oven-grpo-training-data
OVEN GRPO Training Data
Training data behind my MSc thesis GRPO experiments on OVEN open-domain visual
entity recognition with Qwen3-VL-4B. This is the companion to the published
GRPO adapters: the parquets,
taxonomy oracle, and self-aggregation traces they were trained on.
Layout
path
contents
taxonomy/chains_labels.jsonl
Wikidata P279 (subclass-of) chains per OVEN entity: the label hierarchy used by every hierarchical metric… See the full description on the dataset page: https://huggingface.co/datasets/jucamohedano/oven-grpo-training-data.xyran_train_dataset
Xyran training dataset
Prepared for Xyran local content-safety model training.
Layout
sfw/
nsfw/
nsfl/
_manifests/
anime_dbrating WebP migration: 20260905_111217
Mapping:
general + sensitive -> SFW
questionable + explicit -> NSFW
WebP normalization:
existing WebP: byte-for-byte passthrough
non-WebP: libvips -> WebP
quality: 95
lossless: False
effort: 2
no resize
no crop
ZIP payload: STORE
Results:
SFW images: 683,275
NSFW images: 598,226
SFW… See the full description on the dataset page: https://huggingface.co/datasets/MingSafeR/xyran_train_dataset.herislab-ca-training-data
CA_Training_Data -- Convolutional Autoencoder (Track A)
Curated dataset for training and evaluating the Convolutional Autoencoder anomaly detection model.
Approach
The autoencoder is trained only on normal (no-fault) images. At inference, high reconstruction error indicates an anomaly/fault.
Structure
train/normal/ -- Normal images for autoencoder training
electric_motor/ -- 168 PNG (Electric Motor Thermal Fault Diagnosis, no_fault class)… See the full description on the dataset page: https://huggingface.co/datasets/Ryanflash/herislab-ca-training-data.skin-lesion-trainOne-full
Skin Lesion Dataset — trainOne Full
All 7 HAM10000 classes including Melanocytic Nevi (nv).
Source: ISIC 2018 Challenge Task 3.
Code
Description
Count
nv
Melanocytic Nevi (benign)
5000
mel
Melanoma (deadly)
1113
bkl
Benign Keratosis
1099
bcc
Basal Cell Carcinoma (cancer)
514
akiec
Actinic Keratosis (monitor)
327
vasc
Vascular Lesions
142
df
Dermatofibroma
115
Total: 8,310 images
Load:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/zujiguts/skin-lesion-trainOne-full.matt-training-imgMy dataset for training SDXL & SD 1.5
v-lol-trains
Dataset Card for Dataset Name
Dataset Summary
This diagnostic dataset (website, paper) is specifically designed to evaluate the visual logical learning capabilities of machine learning models.
It offers a seamless integration of visual and logical challenges, providing 2D images of complex visual trains,
where the classification is derived from rule-based logic.
The fundamental idea of V-LoL remains to integrate the explicit logical learning tasks of classic symbolic AI… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/v-lol-trains.ai-auto-train-datasets-cuda-5d-quantum-mindmap-simulations-generator-zkevms-immutablexlanternfly_swatter_training
Spotted Lanternfly Classification Dataset
Dataset Description
This dataset contains images for binary classification of spotted lanternflies (Lycorma delicatula), an invasive species causing significant damage to agriculture and ecosystems in the United States. The dataset is designed to train machine learning models to identify dead or squashed lanternflies from photographs, supporting community-driven environmental monitoring and pest management efforts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/rlogh/lanternfly_swatter_training.data-trainfootball-dataset_training
Dataset Description
This dataset contains football (soccer) field images captured from a tactical camera perspective. The dataset contains two images per frame;
One image highlights only the green color, while all other colors are grayed out.
The other image grays out the green color, while all other colors remain in full color.
The dataset is designed for computer vision research in sports analytics, including player tracking, field understanding, and tactical analysis.
Field… See the full description on the dataset page: https://huggingface.co/datasets/chimp-ll/football-dataset_training.Roof_Training_Images_2Experiment-SLAM-6c-custom-training-dataTemporary ad-hoc dataset for use in private training experiment.
tiktok-techjam-2026-openfake-train
TikTok TechJam 2026 — OpenFake train slice
Filtered, re-encoded train split used by Seer. This is not a copy of
ComplexDataLab/OpenFake
(3.44 TB). It is the on-disk mixture source openfake/train.
Do not use core/test or reddit/test from this repo — those holdouts stay
out of training.
Contents
439,523 JPEGs, ~102 GB. Images were streamed from OpenFake core/train,
filtered to 30 generators + two real pools, resized to max_side=1536,
JPEG quality 95.
split… See the full description on the dataset page: https://huggingface.co/datasets/glennwuwu/tiktok-techjam-2026-openfake-train.breast-histopathology-images-train-test-valid-split
Breast Histopathology Image dataset
This dataset is just a rearrangement of the Original dataset at Kaggle: https://www.kaggle.com/datasets/paultimothymooney/breast-histopathology-images
Data Citation: https://www.ncbi.nlm.nih.gov/pubmed/27563488 , http://spie.org/Publications/Proceedings/Paper/10.1117/12.2043872
The original dataset has structure: |-- patient_id
|-- class(0 and 1)
The present dataset has following structure: |-- train
|-- class(0 and 1)
|--… See the full description on the dataset page: https://huggingface.co/datasets/EulerianKnight/breast-histopathology-images-train-test-valid-split.training
87cc2s/training
Numerosity training data, organized by task then by source:
counting-training/
synthetic/ <- abstract shapes generator
real-world/ <- real photographs, ~19 public counting/detection datasets
ans-training/
synthetic/ <- abstract shapes generator
real-world/ <- same-category pairs of real photographs
Each *-training/<domain>/ folder has its own dataset_card.md (schema +
generation/curation details), manifest.json (summary stats), and… See the full description on the dataset page: https://huggingface.co/datasets/87cc2s/training.
