datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.cinic10
Dataset Card for CINIC-10
CINIC-10 has a total of 270,000 images equally split amongst three subsets: train, validate, and test. This means that CINIC-10 has 4.5 times as many samples than CIFAR-10.
Dataset Details
In each subset (90,000 images), there are ten classes (identical to CIFAR-10 classes). There are 9000 images per class per subset. Using the suggested data split (an equal three-way split), CINIC-10 has 1.8 times as many training samples as in CIFAR-10.… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/cinic10.TTA-CIFAR-100-C
TTA-CIFAR-100-C
Mirror of CIFAR-100-C (Hendrycks & Dietterich, ICLR 2019) with a
revision pin for reproducible test-time adaptation evaluation.
Upstream: Zenodo record 3555552
License: CC BY 4.0 (matches upstream)
Sibling: TTA-CIFAR-10-C
Maintained as part of: TTA-Evaluation-Harness
Citation
@inproceedings{hendrycks2019benchmarking,
title={Benchmarking Neural Network Robustness to Common Corruptions and Perturbations},
author={Hendrycks, Dan and Dietterich… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/TTA-CIFAR-100-C.cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar10.CIFAR10
🖼️ CIFAR10 (Extracted from PyTorch Vision)
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
ℹ️ Dataset Details
📖 Dataset Description
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The classes are completely mutually exclusive. There is no… See the full description on the dataset page: https://huggingface.co/datasets/p2pfl/CIFAR10.TTA-CIFAR-10-C
TTA-CIFAR-10-C
Mirror of CIFAR-10-C (Hendrycks & Dietterich, ICLR 2019) with a revision
pin for reproducible test-time adaptation evaluation.
Upstream: Zenodo record 2535967
License: CC BY 4.0 (matches upstream)
SHA256 of upstream tarball: c72763e101c723b7c507b96205f7e938912a5d587376173b825850cf3cb876a7
Maintained as part of: TTA-Evaluation-Harness
Citation
@inproceedings{hendrycks2019benchmarking,
title={Benchmarking Neural Network Robustness to Common… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/TTA-CIFAR-10-C.wine-images-126k
Wine Images Dataset 126K
A comprehensive dataset of 107,821 wine bottle images linked to the Wine Text Dataset 126K. This companion dataset provides high-quality wine bottle images for computer vision, multimodal machine learning, and wine recognition tasks.
Dataset Description
This dataset contains wine bottle images scraped from wine retailer websites. Each image is linked to detailed wine information (descriptions, pricing, categories, regions) via stable IDs that… See the full description on the dataset page: https://huggingface.co/datasets/cipher982/wine-images-126k.cifar100-lt
Dataset Card for CIFAR-100-LT (Long Tail)
Dataset Summary
Note (March 2026): This dataset has been migrated from a Python loading script to
parquet format for compatibility with datasets v4.4+. No trust_remote_code=True
is needed. Available configs: r-10, r-20, r-50, r-100.
The CIFAR-100-LT imbalanced dataset is comprised of under 60,000 color images, each measuring 32x32 pixels,
distributed across 100 distinct classes.
The number of samples within each class… See the full description on the dataset page: https://huggingface.co/datasets/tomas-gajarsky/cifar100-lt.citrus_fruit_variety_classification
Citrus Fruit Variety Classification
A dataset for variety classification of citrus fruits. The dataset contains raw and augmented versions.The raw dataset contains 1,379 images.Images per class:
murcott: 280
ponkan: 328
tangerine: 400
tankan: 371
The augmented dataset contains 7,584 images.Images per class:
murcott: 1,540
ponkan: 1,803
tangerine: 2,200
tankan: 2,041
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/citrus_fruit_variety_classification.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar100.Canadian-streetview-cities
Canadian Street View Cities Dataset
Overview
A street-view image dataset created to train and evaluate models for city-level image classification across major Canadian cities. Each entry includes an image and its corresponding city label.
Purpose
The dataset is intended for building models that recognize the Canadian city in which a street-view scene was captured.
Data Source
All images were collected from Mapillary, using geographic bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/SABR22/Canadian-streetview-cities.citrus-disease-vlm-instruct
Citrus Disease VLM Instruct
An instruction-tuning dataset for training a small vision-language model (VLM) to look at a photo of a citrus leaf, fruit or shoot, name the disease, pest or nutrient deficiency, explain the cause and symptoms, and recommend both biological/organic and chemical management.
Every example pairs one image with a chat conversation in the format used by TRL's SFTTrainer for multimodal models (Qwen-VL, SmolVLM, Idefics, LLaVA and similar).
What… See the full description on the dataset page: https://huggingface.co/datasets/ML-Intern-lab/citrus-disease-vlm-instruct.cifar10-lt
Dataset Card for CIFAR-10-LT (Long Tail)
Dataset Summary
Note (March 2026): This dataset has been migrated from a Python loading script to
parquet format for compatibility with datasets v4.4+. No trust_remote_code=True
is needed. Available configs: r-10, r-20, r-50, r-100.
The CIFAR-10-LT imbalanced dataset is comprised of under 60,000 color images, each measuring 32x32 pixels,
distributed across 10 distinct classes.
The number of samples within each class decreases… See the full description on the dataset page: https://huggingface.co/datasets/tomas-gajarsky/cifar10-lt.naip-16d-city-cubes
NAIP 16-Day City Cubes (materialized tiles)
Each row is a 512×512 chip with 16 layers (composites, single-band indices, and masks).
What’s included (no pseudoRGB)
RGB composites: naip_rgb, s2_rgb, dem_rgb
Mono S2 layers (published as single-channel images): s2_B08, s2_MSAVI, s2_NDVI, s2_NDWI, s2_SCL
Other monos: naip_ndvi
Semantic masks: labels (task labels), landfire_family, cdl
Metadata: tile_id, city, bbox (west,south,east,north), chip_px, split, meta_json
Note:… See the full description on the dataset page: https://huggingface.co/datasets/gdurkin/naip-16d-city-cubes.osm-civic-community-places
OSM Civic & Community Places Visual Dataset
Rows: 397,704
Dataset Description
OSM Civic & Community Places Visual Dataset is a Global wildlife image dataset alternative for geospatial computer vision and image classification focused on OpenStreetMap-mapped civic and community places. The labeled classes are School, Places of Worship, and Nature Reserve, produced by Outerview's query-driven visual matching and label assignment process over OpenStreetMap node… See the full description on the dataset page: https://huggingface.co/datasets/Outerview/osm-civic-community-places.citrusuat_disease_classification
CitrusUAT Disease Classification
A dataset for disease classification of orange leaves. The dataset contains 953 images across 12 classes: Citrus_leafminer, Fe, Greasy_spot, HLB, Healthy, Mg, Mn, N, Red_scale, Red_scale_sequelae, Texas_mite, Zn.Images per class:
Citrus_leafminer: 100
Fe: 100
Greasy_spot: 100
HLB: 43
Healthy: 100
Mg: 100
Mn: 30
N: 50
Red_scale: 30
Red_scale_sequelae: 100
Texas_mite: 100
Zn: 100
This dataset is indexed on https://project-agml.github.io/ as part… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/citrusuat_disease_classification.svae-freckles-4096-cifar10
SVAE Freckles 4096 — CIFAR-10 Omega Tokens
Precomputed spectral decomposition of CIFAR-10 through Freckles v41 (256×256), a frozen Spectral Variational Autoencoder trained exclusively on synthetic noise.
Each CIFAR-10 image is resized to 256×256, decomposed into 4096 patches (4×4 each), and passed through Freckles' encoder → SVD bottleneck. The 4 singular values per patch are stored as a (4, 64, 64) omega map — a 4-channel spatial representation of spectral energy.… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/svae-freckles-4096-cifar10.citrus_fruit_leaf_disease_classification
Citrus Fruit Leaf Disease Classification
A dataset for disease classification of citrus fruits and leaves. The dataset contains 759 images across 6 classes: black_spot, canker, greening, healthy, melanose, scab.Images per class:
black_spot: 190
canker: 241
greening: 220
healthy: 80
melanose: 13
scab: 15
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{rauf2019citrus,
title={A citrus fruits and… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/citrus_fruit_leaf_disease_classification.cifar10h
CIFAR-10H Hugging Face Dataset
This repository contains a Hugging Face dataset build of CIFAR-10H, an extension of the CIFAR-10 test set with human-annotated label distributions.
Dataset Description
CIFAR-10H adds human uncertainty information to the CIFAR-10 test images by providing:
expert_probs: probability distributions over the 10 CIFAR-10 classes
expert_counts: raw human vote counts for each class
expert_argmax: one-hot encoded labels from the human-majority choice… See the full description on the dataset page: https://huggingface.co/datasets/MKZuziak/cifar10h.CIFAR100-customExample of usage:
from datasets import load_dataset
dataset = load_dataset("Andron00e/CIFAR100-custom")
splitted_dataset = dataset["train"].train_test_split(test_size=0.2)
cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may… See the full description on the dataset page: https://huggingface.co/datasets/puneet44/cifar10.cifar10-outlier
Dataset Card for "cifar10-outlier"
📚 This dataset is an enriched version of the CIFAR-10 Dataset.
The workflow is described in the medium article: Changes of Embeddings during Fine-Tuning of Transformers.
Explore the Dataset
The open source data curation tool Renumics Spotlight allows you to explorer this dataset. You can find a Hugging Face Spaces running Spotlight with this dataset here:
Full Version (High hardware requirement)… See the full description on the dataset page: https://huggingface.co/datasets/renumics/cifar10-outlier.cifar100-outlier
Dataset Card for "cifar100-outlier"
📚 This dataset is an enriched version of the CIFAR-100 Dataset.
The workflow is described in the medium article: Changes of Embeddings during Fine-Tuning of Transformers.
Explore the Dataset
The open source data curation tool Renumics Spotlight allows you to explorer this dataset. You can find a Hugging Face Space running Spotlight with this dataset here: https://huggingface.co/spaces/renumics/cifar100-outlier.
Or you can explorer it… See the full description on the dataset page: https://huggingface.co/datasets/renumics/cifar100-outlier.cifar10-lt-federated
CIFAR-10 Long-Tail Federated Dataset
Dataset Description
This is a long-tailed version of CIFAR-10 designed for federated learning research. The dataset introduces class imbalance following an exponential decay distribution, making it ideal for studying long-tail classification in federated settings.
Class Distribution (Training Set)
The training set follows a long-tail distribution with imbalance factor 100:
airplane (Class 0): 5,000 samples
automobile (Class… See the full description on the dataset page: https://huggingface.co/datasets/Beothuk/cifar10-lt-federated.CIFAR-10_Subset
CIFAR-10 — Subset
Stratified random subset of CIFAR-10.
Split
Rows
Per class
train
5,000
500
test
1,000
100
validation
500
50
Classes: airplane, automobile, bird, cat, deer, dog, frog, horse, ship, truck
Images: 32 × 32 RGB | Seed: 42
Label Map
ID
Class
ID
Class
0
airplane
5
dog
1
automobile
6
frog
2
bird
7
horse
3
cat
8
ship
4deer
9
truck
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Chiranjeev007/CIFAR-10_Subset.cifar100-lt
Dataset Card for CIFAR-100-LT (Long Tail)
Dataset Summary
Note (March 2026): This dataset has been migrated from a Python loading script to
parquet format for compatibility with datasets v4.4+. No trust_remote_code=True
is needed. Available configs: r-10, r-20, r-50, r-100.
The CIFAR-100-LT imbalanced dataset is comprised of under 60,000 color images, each measuring 32x32 pixels,
distributed across 100 distinct classes.
The number of samples within each class… See the full description on the dataset page: https://huggingface.co/datasets/Amanmeena004/cifar100-lt.CIFAR10
🖼️ CIFAR10 (Extracted from PyTorch Vision)
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
ℹ️ Dataset Details
📖 Dataset Description
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The classes are completely mutually exclusive. There… See the full description on the dataset page: https://huggingface.co/datasets/rajnandinib/CIFAR10.xai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.CIFAR10-customExample of usage:
from datasets import load_dataset
dataset = load_dataset("Andron00e/CIFAR10-custom")
splitted_dataset = dataset["train"].train_test_split(test_size=0.2)
