datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.AID
Aerial Image Dataset (AID)
Description
The Aerial Image Dataset (AID) is a scene classification dataset consisting of 10,000 RGB images, each with a resolution of 600x600 pixels. These images have been extracted using Google Earth and cover various scenes from regions and countries around the world. AID comprises 30 different scene categories, with several hundred images per class.
The new dataset is made up of the following 30 aerial scene types: airport, bare… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/AID.AIGIBench
Is Artificial Intelligence Generated Image Detection a Solved Problem?
Ziqiang Li1, Jiazhen Yan1, Ziwen He1, Kai Zeng2, Weiwei Jiang1, Lizhi Xiong1, Zhangjie Fu1‡
‡Corresponding author
1Nanjing University of Information Science and Technology 2University of Siena
Paper | GitHub Repository
This repository is the official dataset of the AIGIBench.
AIGIBench dataset contains two types of training and 25 test subsets. This dataset has the following advantages:
Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/HorizonTEL/AIGIBench.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
The dataset was initially developed exclusively for military aircraft detection, but was later expanded to include commercial airliners for a broader and more challenging detection task.
The dataset contains 103 military aircraft types and 11 commercial airliner types.
Military aircraft: A10, A400M, AG600, AH64, AKINCI, AV8B, An124, An22, An225, An72, B1, B2, B21, B52, Be200, C1… See the full description on the dataset page: https://huggingface.co/datasets/a2015003713/military-aircraft-detection-dataset.brain-structureA collection of T1-weighted .nii.gz structural MRI scans in a BIDS-like arrangement,
with JSON sidecar metadata indicating train/validation/test splits.stable-diffusion-v1-5-glazed
Dataset Card for Stable Diffusion v1.5 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by runwayml/stable-diffusion-v1-5
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/stable-diffusion-v1-5-glazed.FGVC-Aircraft
Dataset Card for FGVC-Aircraft
This is a FiftyOne dataset with 10000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("Voxel51/FGVC-Aircraft")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FGVC-Aircraft.anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.Kuro-Siwo-GeoTIFFs
Kuro Siwo GeoTIFFs
Paper | GitHub |
Dataset Details
Dataset Description
Kuro Siwo is a global multi-temporal SAR dataset for rapid flood mapping. It contains 43 flood events in 6 continents and 3 climate zones, over the period 2015-2022. The annotations have been produced through meticulous photointerpretation by a team of experts, at 10m spatial resolution. For each flood event, we provide one Sentinel-1 post-flood and two Sentinel-1 pre-flood… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Kuro-Siwo-GeoTIFFs.real-fake-ai-generated-art-images
🎨 Real and Fake (AI-Generated) Art Images Dataset
21,642 balanced images — 10,821 real artworks and 10,821 AI-generated
images — for training models to distinguish authentic art from GAN-generated fakes.
🧭 Overview
This dataset is part of the FauxFinder project, designed to build
advanced models capable of distinguishing between authentic artworks
and AI-generated images. Ideal for binary classification, GAN research,
and computer vision benchmarking.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/real-fake-ai-generated-art-images.AIGC-Detection-Benchmark
AIGC Detection Benchmark Dataset
📝 Dataset Description
Dataset Summary
The AIGC Detection Benchmark Dataset is a high-quality collection of images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. The dataset contains a mix of real-world images and images generated by a wide array of prominent AI models, including diffusion models (like Stable Diffusion, DALL-E 2, Midjourney, ADM) and GANs… See the full description on the dataset page: https://huggingface.co/datasets/TheKernel01/AIGC-Detection-Benchmark.visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.Kuro-Siwo-Webdataset
Kuro Siwo webdatasets
Paper | GitHub |
Dataset Details
Dataset Description
Kuro Siwo is a global multi-temporal SAR dataset for rapid flood mapping. It contains 43 flood events in 6 continents and 3 climate zones, over the period 2015-2022. The annotations have been produced through meticulous photointerpretation by a team of experts, at 10m spatial resolution. For each flood event, we provide one Sentinel-1 post-flood and two Sentinel-1 pre-flood… See the full description on the dataset page: https://huggingface.co/datasets/orion-ai-lab/Kuro-Siwo-Webdataset.AI_REAL
AI vs REAL Image Dataset
Ce dataset contient deux classes d’images :
AI : images générées par intelligence artificielle
REAL : images réelles
Les fichiers sont organisés par batchs pour respecter les limites de Hugging Face.
AI-GenBench-fake_part
AI-GenBench: A New Ongoing Benchmark for AI-Generated Image Detection
Important: this is the fake part of the AI-GenBench dataset. To re-create the original benchmark, which includes real images, please check the official repository.
Important 2: before using, please check the licensing terms of the images included!
Details
The rapid advancement of generative AI has revolutionized image creation, enabling high-quality synthesis from text prompts while raising critical… See the full description on the dataset page: https://huggingface.co/datasets/lrzpellegrini/AI-GenBench-fake_part.cifar-10-python
CIFAR-10, original python archive
Unmodified copy of cifar-10-python.tar.gz from https://www.cs.toronto.edu/~kriz/cifar.html,
mirrored so course notebooks do not depend on the original host.
MD5 c58f30108f718f92721af3b95e74349a (the value torchvision.datasets.CIFAR10 checks)
170,498,071 bytes
Use with torchvision
Set the URL before the first datasets.CIFAR10(...) call. torchvision still verifies the
MD5, so nothing else changes. Pin a commit hash from this… See the full description on the dataset page: https://huggingface.co/datasets/MIT-OL-AI-D/cifar-10-python.RAIDThis dataset is for testing the adversarial robustness of AI-Generated Image Detectors, as described in the paper RAID: A Dataset for Testing the Adversarial Robustness of AI-Generated Image Detectors.
Military-Aircraft-DetectionDataset for object detection of military aircraft
bounding box in PASCAL VOC format (xmin, ymin, xmax, ymax)
43 aircraft types
(A-10, A-400M, AG-600, AV-8B, B-1, B-2, B-52 Be-200, C-130, C-17, C-2, C-5, E-2, E-7, EF-2000, F-117, F-14, F-15, F-16, F/A-18, F-22, F-35, F-4, J-20, JAS-39, MQ-9, Mig-31, Mirage2000, P-3(CP-140), RQ-4, Rafale, SR-71(may contain A-12), Su-34, Su-57, Tornado, Tu-160, Tu-95(Tu-142), U-2, US-2(US-1A Kai), V-22, Vulcan, XB-70, YF-23)
Please let me know if you find wrong… See the full description on the dataset page: https://huggingface.co/datasets/Illia56/Military-Aircraft-Detection.ai-image-detector-dataset
AI Image Detector Dataset (training-v1)
wkaandemir/ai-image-detector modelini eğitmek için kullanılan, 20.000 normalize edilmiş görselden oluşan dengeli ve kaynak-farkında (source-aware) bir görüntü sınıflandırma veri kümesi. Görev, görselleri gerçek (real) ve yapay (fake) olarak ikili sınıflandırmaktır.
Bu veri kümesi, modelin kalibrasyon ve eşik seçiminde kullanılmayan, genelleme ölçümü için kaynak bazlı ayrı tutulmuş bir external_test split'i de içerir.
📌… See the full description on the dataset page: https://huggingface.co/datasets/wkaandemir/ai-image-detector-dataset.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.PolyMFO
PolyMFO
PolyMFO is a high-resolution image dataset for foreign-object anomaly classification and binary segmentation. It contains normal samples and samples from nine foreign-object classes. Every image has a corresponding binary mask.
This Hugging Face repository contains the dataset only.Source code for data processing, benchmark preparation, evaluation, and related research utilities is maintained at https://github.com/CaoMinhh/PolyMFO.
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/VNSO-AI/PolyMFO.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
For the newest updates on this dataset and our related research, see
https://scam.ai/research.
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.military-aircraft-detection-dataset
Military Aircraft Detection Dataset
Military aircraft detection dataset in COCO and YOLO format.
This dataset is synchronized from the original Kaggle dataset:https://www.kaggle.com/datasets/a2015003713/militaryaircraftdetectiondataset
aidovecl-vehicle-detection-classification-localization
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
We introduce an annotated AI-generated dataset of eye-level vehicle images using outpainting, offering versatile generation of diverse vehicle classes in varied contexts with pretrained models.
Citation Notice
Please ensure that all publications and presentations using this data reference the following paper:
Kazemi, A., Fatima, Q. ul A., Kindratenko, V., & Tessum, C. W.… See the full description on the dataset page: https://huggingface.co/datasets/amir-kazemi/aidovecl-vehicle-detection-classification-localization.aidm-dogs-vs-cats-data
aidm-dogs-vs-cats-data
The images of a dogs-vs-cats image-classification study, in the directory layout that results/splits.json indexes. The run registry, the splits file, the report tables and figures live in the companion results repo; the checkpoints live in the companion model repo.
Provenance
These images are the Kaggle Dogs vs. Cats competition data (https://www.kaggle.com/c/dogs-vs-cats). They are not original work and they are not relicensed here. They… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/aidm-dogs-vs-cats-data.aircraft-images
Dataset Card for High-Resolution Aircraft Images
Dataset Summary
This dataset contains 165,340 high-resolution aircraft images collected from the internet, along with machine-generated captions. The captions were generated using Gemini Flash 1.5 AI model and are stored in separate text files matching the image filenames.
Languages
The dataset is monolingual:
English (en): All image captions are in English
Dataset Structure
Data Files
The… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/aircraft-images.aidRedistributed without modification from
https://huggingface.co/datasets/blanchon/AID, itself a rehost of Xia et al.
2017's Aerial Image Dataset. 10,000 RGB images, 600x600px, 30 scene classes,
extracted from Google Earth.
License is unspecified upstream -- neither the original paper nor the
blanchon/AID rehost states one. Verify suitability for your use case before
redistributing further.
Please cite https://doi.org/10.1109/TGRS.2017.2685945 if you use this dataset.
The train/val/test split… See the full description on the dataset page: https://huggingface.co/datasets/isaaccorley/aid.aidetector-data
Companion dataset for the aidetector
project (live demo: https://humanorai.online). Code, trained model, and the
reproducibility record live in that GitHub repo.
Dataset Card — AI Image Detector
This card documents the data the v2 models were trained and evaluated on. All
counts are taken from the reproducibility record
experiment_v1.json (dataset content hash
72b88efc0497...). The image files themselves are not redistributed in this
repository (see Access below).
The intent… See the full description on the dataset page: https://huggingface.co/datasets/aman213/aidetector-data.
