datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar10.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/uoft-cs/cifar100.cinic10
Dataset Card for CINIC-10
CINIC-10 has a total of 270,000 images equally split amongst three subsets: train, validate, and test. This means that CINIC-10 has 4.5 times as many samples than CIFAR-10.
Dataset Details
In each subset (90,000 images), there are ten classes (identical to CIFAR-10 classes). There are 9000 images per class per subset. Using the suggested data split (an equal three-way split), CINIC-10 has 1.8 times as many training samples as in CIFAR-10.… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/cinic10.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/wallstoneai/civitai-top-nsfw-images-with-metadata.cifar100-enrichedThe CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).cilp_assessment_subset
Dataset Card for cilp_assessment_subset
Created with this Jupyter Notebook.
This is a FiftyOne dataset with 750 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("philippkolbe/cilp_assessment_subset")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/philippkolbe/cilp_assessment_subset.TTA-CIFAR-100-C
TTA-CIFAR-100-C
Mirror of CIFAR-100-C (Hendrycks & Dietterich, ICLR 2019) with a
revision pin for reproducible test-time adaptation evaluation.
Upstream: Zenodo record 3555552
License: CC BY 4.0 (matches upstream)
Sibling: TTA-CIFAR-10-C
Maintained as part of: TTA-Evaluation-Harness
Citation
@inproceedings{hendrycks2019benchmarking,
title={Benchmarking Neural Network Robustness to Common Corruptions and Perturbations},
author={Hendrycks, Dan and Dietterich… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/TTA-CIFAR-100-C.City-Landscape-In-Sight
City Landscape In Sight — Window View Perception (Images & Models)
This dataset hosts the large binary artefacts for the paper "City landscape in sight: A crowdsourced framework for unlocking urban-scale window view perceptions from real estate imagery." It is the companion of the code repository on GitHub:
Paper (arXiv): https://arxiv.org/abs/2606.15198
Code & derived data (GitHub): https://github.com/Sijie-Yang/City-Landscape-In-Sight
Images & trained weights (this dataset):… See the full description on the dataset page: https://huggingface.co/datasets/sijiey/City-Landscape-In-Sight.cifar10
Dataset Card for CIFAR-10
Dataset Summary
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
The dataset is divided into five training batches and one test batch, each with 10000 images. The test batch contains exactly 1000 randomly-selected images from each class. The training batches contain the remaining images in random order, but some training batches may contain… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar10.CIFAR10
🖼️ CIFAR10 (Extracted from PyTorch Vision)
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images.
ℹ️ Dataset Details
📖 Dataset Description
The CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images per class. There are 50000 training images and 10000 test images. The classes are completely mutually exclusive. There is no… See the full description on the dataset page: https://huggingface.co/datasets/p2pfl/CIFAR10.TTA-CIFAR-10-C
TTA-CIFAR-10-C
Mirror of CIFAR-10-C (Hendrycks & Dietterich, ICLR 2019) with a revision
pin for reproducible test-time adaptation evaluation.
Upstream: Zenodo record 2535967
License: CC BY 4.0 (matches upstream)
SHA256 of upstream tarball: c72763e101c723b7c507b96205f7e938912a5d587376173b825850cf3cb876a7
Maintained as part of: TTA-Evaluation-Harness
Citation
@inproceedings{hendrycks2019benchmarking,
title={Benchmarking Neural Network Robustness to Common… See the full description on the dataset page: https://huggingface.co/datasets/WNJXYK/TTA-CIFAR-10-C.cardiac_cine_acdc
ACDC (Cardiac Cine-MRI)
ACDC (Automatic Cardiac Diagnosis Challenge, MICCAI 2017) is a cine‑MRI dataset for cardiac segmentation.This repository contains processed NIfTI files in Data/processed_output/acdc format.
Dataset Summary
Modality: Cardiac cine‑MRI (NIfTI)
Task: Segmentation of LV, RV, and myocardium
Frames: ED/ES + full SAX time series (sax_t)
Labels: LV/RV cavities + myocardium
Splits: train, test (as provided in processed output)
Data Structure (per… See the full description on the dataset page: https://huggingface.co/datasets/viennh2012/cardiac_cine_acdc.cifar10
Dataset Specifications
Contains the entire CIFAR10 dataset, downloaded via PyTorch, then split and saved as .png files representing 32x32 images.
There a three splits, perfectly balanced class-wise:
train: 49,000 out of the original 50,000 samples from the training set of CIFAR10;
calibration: 1,000 left-out samples from the training set;
test: 10,000 samples, the entire original test set.
File Structure
Files are archives <split>/<classname>.zip. Each… See the full description on the dataset page: https://huggingface.co/datasets/ego-thales/cifar10.aigriculture-challenge-2026-track1-plantlet-viability
AI-griculture Challenge 2026 · Track 1: AI-based plantlet viability assessment
Photographs of in-vitro potato plantlets from the genebank of the International Potato Center
(CIP, Lima, Peru), taken during routine viability monitoring. Each photo shows the test tubes
(one to three) of one genebank accession. The task is to classify each photo into one of four
viability classes: Good, Medium, Regular, Critical. The class of a photo is the viability
state of the accession, judged… See the full description on the dataset page: https://huggingface.co/datasets/cipotato/aigriculture-challenge-2026-track1-plantlet-viability.wine-images-126k
Wine Images Dataset 126K
A comprehensive dataset of 107,821 wine bottle images linked to the Wine Text Dataset 126K. This companion dataset provides high-quality wine bottle images for computer vision, multimodal machine learning, and wine recognition tasks.
Dataset Description
This dataset contains wine bottle images scraped from wine retailer websites. Each image is linked to detailed wine information (descriptions, pricing, categories, regions) via stable IDs that… See the full description on the dataset page: https://huggingface.co/datasets/cipher982/wine-images-126k.CIFAR-CThe license is to the original authors (see below)!
This repository contains the CIFAR-10-C dataset from Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. We are currently hosting it on Hugging Face due to an increased latency from Zenodo.
We are not the original authors. If you find this useful in your research, please consider citing:
@article{hendrycks2019robustness,
title={Benchmarking Neural Network Robustness to Common Corruptions and Perturbations}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/CIFAR-C.cifar10-enrichedThe CIFAR-10 dataset consists of 60000 32x32 colour images in 10 classes, with 6000 images
per class. There are 50000 training images and 10000 test images.
This version if CIFAR-10 is enriched with several metadata such as embeddings, baseline results and label error scores.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/Avanish11/civitai-top-nsfw-images-with-metadata.MICO-CIFAR10
MICO CIFAR-10 challenge dataset
Mico Argentatus (Silvery Marmoset) - William Warby/Flickr
For the accompanying code, visit the GitHub repository of the competition: https://github.com/microsoft/MICO/.
Getting Started
The starting kit notebook for this task is available at: https://github.com/microsoft/MICO/tree/main/starting-kit.
In the starting kit notebook you will find a walk-through of how to load the data and make your first submission.
We also provide a… See the full description on the dataset page: https://huggingface.co/datasets/szanella/MICO-CIFAR10.BTL3_CIFAR-10
CIFAR-10 Feature Representations (BTL3)
This dataset contains pre-extracted feature embeddings from the CIFAR-10 dataset, produced using several pretrained image classification models.The goal is to enable fast experimentation, classifier prototyping, and model comparison without needing to train or forward pass large models in Colab.
Dataset Source
The original CIFAR-10 dataset is MIT-licensed and available here:https://www.cs.toronto.edu/~kriz/cifar.html
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LeTienDat/BTL3_CIFAR-10.cifar100-lt
Dataset Card for CIFAR-100-LT (Long Tail)
Dataset Summary
Note (March 2026): This dataset has been migrated from a Python loading script to
parquet format for compatibility with datasets v4.4+. No trust_remote_code=True
is needed. Available configs: r-10, r-20, r-50, r-100.
The CIFAR-100-LT imbalanced dataset is comprised of under 60,000 color images, each measuring 32x32 pixels,
distributed across 100 distinct classes.
The number of samples within each class… See the full description on the dataset page: https://huggingface.co/datasets/tomas-gajarsky/cifar100-lt.temporal-aerial-cityline-construction-sample
CityLine — Temporal Aerial Construction Dataset (Sample)
Temporal Aerial Vision · Construction Progress · Multiview Geometry · San Jose, CA
CityLine is a multi-year aerial imagery sequence captured from a helicopter during the construction of a major mixed-use development in San Jose, California.This sample highlights multiple construction phases over time, with several oblique views per capture date.
The full (commercial) dataset contains hundreds of high-resolution images with… See the full description on the dataset page: https://huggingface.co/datasets/SharpShots/temporal-aerial-cityline-construction-sample.civitai-top-nsfw-images-with-metadata
CivitAI Top NSFW Images Dataset
This dataset contains 6k+ top NSFW images from CivitAI filtered using top reactions. The dataset contains prompt & nsfw level metadata in prompts.json file. The nsfw levels are: Soft, Mature & X.
Original forum post:
https://diffused.to/Thread-CivitAI-Top-NSFW-Images-Dataset-6k-images
Dataset collection date
June 2025
Dataset structure:
├── 📂 images/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ ├── ....
├──… See the full description on the dataset page: https://huggingface.co/datasets/pasindu29/civitai-top-nsfw-images-with-metadata.cifar-10-python
CIFAR-10, original python archive
Unmodified copy of cifar-10-python.tar.gz from https://www.cs.toronto.edu/~kriz/cifar.html,
mirrored so course notebooks do not depend on the original host.
MD5 c58f30108f718f92721af3b95e74349a (the value torchvision.datasets.CIFAR10 checks)
170,498,071 bytes
Use with torchvision
Set the URL before the first datasets.CIFAR10(...) call. torchvision still verifies the
MD5, so nothing else changes. Pin a commit hash from this… See the full description on the dataset page: https://huggingface.co/datasets/MIT-OL-AI-D/cifar-10-python.ciln-bench-cifar10
CILN-Bench: CIFAR-10
This is a label noise dataset. We corrupt the CIFAR-10 noisy-label-train split, 22,500 images, with a known corruption type and severity. Four trained classifiers, called voters, label the corrupted inputs. Each voter's top class is one vote. The vote shares are the label distribution of the input. The true labels, every voter's softmax and the corruption seeds are included, so you can build any kind of label you want.
Settings
15 corruption… See the full description on the dataset page: https://huggingface.co/datasets/sh-islam/ciln-bench-cifar10.cifar10_quality_driftThis dataset was crafted to be used in our tutorial [Link to the tutorial when
ready]. It consists on product reviews from an e-commerce store. The reviews
are labeled on a scale from 1 to 5 (stars). The training & validation sets are
fully composed by reviews written in english. However, the production set has
some reviews written in spanish. At Arize, we work to surface this issue and
help you solve it.citrus_fruit_variety_classification
Citrus Fruit Variety Classification
A dataset for variety classification of citrus fruits. The dataset contains raw and augmented versions.The raw dataset contains 1,379 images.Images per class:
murcott: 280
ponkan: 328
tangerine: 400
tankan: 371
The augmented dataset contains 7,584 images.Images per class:
murcott: 1,540
ponkan: 1,803
tangerine: 2,200
tankan: 2,041
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
The original… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/citrus_fruit_variety_classification.cifar100
Dataset Card for CIFAR-100
Dataset Summary
The CIFAR-100 dataset consists of 60000 32x32 colour images in 100 classes, with 600 images
per class. There are 500 training images and 100 testing images per class. There are 50000 training images and 10000 test images. The 100 classes are grouped into 20 superclasses.
There are two labels per image - fine label (actual class) and coarse label (superclass).
Supported Tasks and Leaderboards
image-classification: The… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cifar100.CIFAKE_autotrain_compatible
Dataset Card for CIFAKE_autotrain_compatible
Dataset Summary
This is a copy of the CIFAKE dataset created by Dr Jordan J. Bird and Professor Ahmad Lotfi. See more information on the original data card on Kaggle.
The real images used are from CIFAR-10. The fake images were created by the authors using Stable Diffusion v1.4.
This dataset removes the train/test structures in the original dataset to allow compatibility with HuggingFace's AutoTrain. It removes the test split… See the full description on the dataset page: https://huggingface.co/datasets/yanbax/CIFAKE_autotrain_compatible.Canadian-streetview-cities
Canadian Street View Cities Dataset
Overview
A street-view image dataset created to train and evaluate models for city-level image classification across major Canadian cities. Each entry includes an image and its corresponding city label.
Purpose
The dataset is intended for building models that recognize the Canadian city in which a street-view scene was captured.
Data Source
All images were collected from Mapillary, using geographic bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/SABR22/Canadian-streetview-cities.
