datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset was initially developed exclusively for military aircraft detection, but was later expanded to include commercial airliners for a broader and more challenging detection task.
The dataset contains 103 military aircraft types and 11 commercial airliner types.
Military aircraft: A10, A400M, AG600, AH64, AKINCI, AV8B, An124, An22, An225, An72, B1, B2, B21, B52, Be200, C1… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.Dataset
MM-OphBench: Multi-Center Multimodal Clinical Ophthalmic Benchmark Dataset
A Large-Scale, Standardized Multi-Center Benchmark Covering 7 Imaging Modalities & 4.3M+ Clinical Records
1. Executive Summary & Repository Overview
The MM-OphBench repository hosts a petabyte-scale, clinically harmonized ophthalmic image archive compiled from leading ophthalmic hospitals and benchmark cohorts. It spans 4,307,415 high-resolution diagnostic images and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Kaphathy/Dataset.Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.cctv-datasets
CCTV Datasets for helmet detection + ANPR
Training and evaluation data used by vivekvar/helmet-v5 and vivekvar/helmet-v4.
Source: Andhra Pradesh RTGS CCTV feeds (public road cameras). All crops and frames are from motorcycle traffic scenes.
Folders
Folder
Contents
Purpose
merged_v3/
YOLO-format dataset (data.yaml + train/valid/test)
Bike + rider detection training
clean_merged_data/
Cleaned / deduped crop set
Base training data for v4
extra_khadatkar/… See the full description on the dataset page: https://huggingface.co/datasets/vivekvar/cctv-datasets.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.llbench-dataset
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences
Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute.
LL-Bench is a large-scale, human-preference benchmark for evaluating low-level
vision restoration in the era of large generative models (LGMs). It compares
10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.toothbrush-v2-dataset
Toothbrushing Detection Dataset (v2)
Video and image data for detecting toothbrushing behavior, collected for a Raspberry Pi Zero 2W toothbrush-detection project (toothbrush_v2). A single-class object detector is trained on this data to output [x, y, w, h, confidence] for the toothbrush in frame.
Dataset structure
Files are packed into tar shards (rather than uploaded individually) to stay within the Hub's per-repo file-count guidelines. To reconstruct the… See the full description on the dataset page: https://huggingface.co/datasets/Schrodin-purrrrr/toothbrush-v2-dataset.qev-data
QEV data
Historical training snapshots for QEV, the LAYA-inspired Qwen3.5-2B decision model.
Model.
Configurations overlap. Do not concatenate them or assume independent test sets.
ZIPs in corpora/ use qev-<stage>.zip and contain eligible
original records, images and license notices. Extracted stage folders and original
record/source IDs preserve the recorded training provenance.
Viewer rows expose request/target schemas as JSON strings; parse with json.loads.
IDs, group IDs… See the full description on the dataset page: https://huggingface.co/datasets/ken-jo/qev-data.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.Icarus-dataset
Icarus
A unified multi-modal curriculum dataset for evolutionary neural architecture search. Every row is one self-contained Task = {meta, support, query}, where support and query are lists of (input_Field, output_Field) pairs. The inner loop trains on support; fitness is scored on query. Support is non-empty for every task. Encoders read the Field descriptor (axes, value_type, n_classes, value_range, mask); mask is True where a value is padding/ignored. meta.class_names, when… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/Icarus-dataset.Comiman-Dataset
Comiman Dataset
Attribution is required for every use: Comiman Dataset by Waheed (huggingface.co/Waheed786dar) - https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset
A license-gated comics and manga page corpus built for training a model that can plan and draw full
comic/manga series (the planned model: Waheed786dar/Comiman).
Every book passed an automatic license gate (Creative Commons / CC0 / Public Domain Mark metadata, or a
public-domain claim limited to works… See the full description on the dataset page: https://huggingface.co/datasets/Waheed786dar/Comiman-Dataset.carla-autopilot-multimodal-dataset
CARLA Autopilot Multimodal Dataset
This dataset contains synchronized multimodal driving data collected in the CARLA simulator using the autopilot feature. It provides RGB images from multiple cameras, semantic segmentation, LiDAR point clouds, 2D bounding boxes, and ego-vehicle state/control signals across varied weather, maps, and traffic densities.
The dataset is designed for research in autonomous driving, sensor fusion, imitation learning, and self-driving evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/immanuelpeter/carla-autopilot-multimodal-dataset.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.marine-animals-multimodal-dataset
Marine Animals Multimodal Dataset 🐋
A comprehensive multimodal dataset combining audio recordings and images of 32 marine species.
Dataset Summary
Total samples: 24,911
Species: 32
Audio files: 1,357 unique recordings
Images: 581 (309 matched + 272 from iNaturalist)
Features
species (string): Species name
label (int32): Numeric label (0–31)
audio (Audio): Audio recording of the species
image (Image): Species image
image_index (int32): Image number… See the full description on the dataset page: https://huggingface.co/datasets/Hariprasath5128/marine-animals-multimodal-dataset.700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
This dataset is synchronized from the original Kaggle dataset:https://www.kaggle.com/datasets/a2015003713/militaryaircraftdetectiondataset
PUUM-koa-restoration-camera-trap-dataset
Dataset Card for Koa Associated Biodiversity Camera Trap Dataset
This dataset is aimed at classification of birds visiting planted Acacia koa (koa) trees in the Pu'u Maka'ala Natural Area Reserve (PUUM) on the island of Hawaii (Big Island). The dataset contains full and cropped images collected by camera trap. These images were collected from January 24th to February 25th, 2025.
Dataset Details
This dataset is aimed at classification of birds visiting planted Acacia… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/PUUM-koa-restoration-camera-trap-dataset.plism-dataset-tiles
PLISM dataset
The Pathology Images of Scanners and Mobilephones (PLISM) dataset was created by (Ochi et al., 2024) for the evaluation of AI models’ robustness to inter-institutional domain shifts.
All histopathological specimens used in creating the PLISM dataset were sourced from patients who were diagnosed and underwent surgery at the University of Tokyo Hospital between 1955 and 2018.
PLISM-wsi consists in a group of consecutive slides digitized under 7 different scanners and… See the full description on the dataset page: https://huggingface.co/datasets/owkin/plism-dataset-tiles.liquidrandom-data
liquidrandom-data
Diverse seed data for ML/LLM training data generation pipelines.
Used by the liquidrandom Python package.
Dataset Summary
This dataset contains 520,080 seed data samples across 24 categories,
generated using a hierarchical taxonomy tree approach with LLM-based quality validation
and fuzzy deduplication. Data is stored as Parquet with zstd compression.
Categories
Category
Samples
File
Coding Tasks
30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.DataFontID
DataFontID
A stratified synthetic corpus of 1,199,562 zero-margin text strips for multi-attribute typographic recognition. Every image carries four labels: font family, language, text color, and typographic style.
Images
1,199,562 (1,079,605 train / 59,978 val / 59,979 test)
Font families
75
Languages
11 across Latin, Cyrillic, Arabic, and Han scripts
Colors
64 (EGA palette)
Styles
Regular, Bold, Italic, Bold-Italic
Backgrounds
Texture, Noise, Complex… See the full description on the dataset page: https://huggingface.co/datasets/issai/DataFontID.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.trade_vision_dataset
TradeVision: Hierarchical Physical Business & Multimodal Retail Provenance Dataset
This dataset is continuously seeded from OpenStreetMap, matched to Google Place IDs, harvested for temporal store photos, and enriched with zero-shot computer vision using Hugging Face Hub native pipelines.
Dataset Structure
The dataset is partitioned into three relational subsets loadable via Hugging Face datasets:
from datasets import load_dataset
# 1. Load Canonical Businesses… See the full description on the dataset page: https://huggingface.co/datasets/drksci/trade_vision_dataset.cub200_dataset
Dataset Card for CUB_200_2011
Dataset Summary
The Caltech-UCSD Birds 200-2011 dataset (CUB-200-2011) is an extended version of the original CUB-200 dataset, featuring photos of 200 bird species primarily from North America. This 2011 version significantly expands its predecessor by doubling the number of images per class and introducing new part location annotations, alongside collecting detailed natural language descriptions for each image through Amazon Mechanical Turk… See the full description on the dataset page: https://huggingface.co/datasets/cassiekang/cub200_dataset.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.handwritten-digit-dataset
Handwritten Digit Dataset
This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability.
Dataset Details
The images are pre-processed to match the standard machine learning format for digit recognition:
Dimensions: 28x28 pixels.
Format: Grayscale (single channel).
Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.
