datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
TreeOfLife-10M
Dataset Card for TreeOfLife-10M
Dataset Summary
With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M.imagenet-12k-wds
Dataset Summary
This is a filtered copy of the full ImageNet dataset consisting of the top 11821 (of 21841) classes by number of samples. It has been used to pretrain a number of in12k models in timm.
The code and metadata for building this dataset from the original full ImageNet can be found at https://github.com/rwightman/imagenet-12k
NOTE: This subset was filtered from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-12k-wds.dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/imagenet-1k-wds.imagenet-22k-wds
Dataset Summary
This is a copy of the full ImageNet dataset consisting of all of the original 21841 clases. It also contains labels in a separate field for the '12k' subset described at at (https://github.com/rwightman/imagenet-12k, https://huggingface.co/datasets/timm/imagenet-12k-wds)
This dataset is from the original fall11 ImageNet release which has been replaced by the winter21 release which removes close to 3000 synsets containing people, a number of these are of an offensive… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-22k-wds.imagenet-1k-wds
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-1k-wds.glint360k-wds-gz
Glint360K
This dataset is introduced in the Partial FC paper https://arxiv.org/abs/2010.05222.
There are 17,091,657 images and 360,232 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_. The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 1,385 shards in total.… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/glint360k-wds-gz.Glint360k
Dataset Card for Glint360K
Citiation by InsightFace Repository
We clean, merge, and release the largest and cleanest face recognition dataset Glint360K, which contains 17091657 images of 360232 individuals. By employing the Patial FC training strategy, baseline models trained on Glint360K can easily achieve state-of-the-art performance. Detailed evaluation results on the large-scale test set (e.g. IFRT, IJB-C and Megaface) are as follows:
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yayoimizuha/Glint360k.Semi-Truths
Semi Truths Dataset: A Large-Scale Dataset for Testing Robustness of AI-Generated Image Detectors (NeurIPS 2024 Track Datasets & Benchmarks Track)
Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation?
To address these questions, we introduce Semi-Truths, featuring 27, 600 real images, 223, 400 masks, and 1, 472, 700… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths.Quantum
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/Miku26727/Quantum.pxhere
Dataset Card for pxhere Images
Dataset Summary
This dataset contains a large collection of high-quality photographs sourced from pxhere.com, a free stock photo website. The dataset includes approximately 1,100,000 images in full resolution covering a wide range of subjects including nature, people, urban environments, objects, animals, and landscapes. All images are provided under the Creative Commons Zero (CC0) license, making them freely available for personal and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/pxhere.pexels-tagger-v0-w640-ws-full
Pexels Tagger V0 Webdataset Full Dataset
This is the webdataset dataset for animetimm/pexels-wdtagger-w640.
Images here are resized to min(width, height) <= 640.
How to Use It
from datasets import load_dataset
dataset = load_dataset('animetimm/pexels-tagger-v0-w640-ws-full')
print(dataset["train"][0])
Images
3122908 images in total.
Split
Image Count
Total Size
train
2810634
184 GB
test
156409
10.2 GB
val
155865
10.2 GB
Tags… See the full description on the dataset page: https://huggingface.co/datasets/animetimm/pexels-tagger-v0-w640-ws-full.imagenet-w21-webp-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with WEBP encoded images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
This is the same contents as https://huggingface.co/datasets/timm/imagenet-w21-wds but encoded in webp at ~56% of the size, shard count… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-webp-wds.forbin_dataset
Forbin Dataset: A collection of historical photographs with archival metadata
This repository hosts the Forbin Dataset, a large-scale collection of historical photographs taken or collected by Victor Forbin (1868–1947).
This HuggingFace dataset version provides:
COCO-style annotations (segmentation polygons)
Archival metadata (Box ID, description, notes, dates when available)
A lightweight explorer interface (HTML/JS) to preview images and annotations:… See the full description on the dataset page: https://huggingface.co/datasets/mchelali/forbin_dataset.ms1mv3-wds
MS-Celeb-1M (v3)
This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.
There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was downloaded from https://github.com/deepinsight/insightface/tree/master/recognition/_datasets_ (MS1M-RetinaFace). The original dataset format is MXNet RecordIO. It was converted to WebDataset in this copy here. There are 100… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds.CleanSTL-10
Dataset Card for STL-10 Cleaned (Deduplicated Training Set)
Paper | Code
Dataset Description
This dataset is a modified version of the STL-10 dataset. The primary modification involves deduplicating the training set by removing any images that are exact byte-for-byte matches (based on SHA256 hash) with images present in the original STL-10 test set. The dataset comprises this cleaned training set and the original, unmodified STL-10 test set.
The goal is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Shu1L0n9/CleanSTL-10.SPIDER-skin
SPIDER-SKIN Dataset
SPIDER is a collection of supervised pathological datasets covering multiple organs, each with comprehensive class coverage. These datasets are professionally annotated by pathologists.
If you would like to support, sponsor, or obtain a commercial license for the SPIDER data and models, please contact us at models@hist.ai.
For a detailed description of SPIDER, methodology, and benchmark results, refer to our research paper:
SPIDER: A Comprehensive Multi-Organ… See the full description on the dataset page: https://huggingface.co/datasets/histai/SPIDER-skin.ms1mv3-wds-gz
MS-Celeb-1M (v3)
This is a copy of gaunernst/ms1mv3-wds with gzip compression. Thus, the shards (after decompression) are identical. The compression ratio is around 50%, indicating that the original JPEG images were not compressed much.
This dataset is introduced in the Lightweight Face Recognition Challenge at ICCV 2019. Paper.
There are 5,179,510 images and 93,431 ids. All images are aligned based on facial landmarks predicted by RetinaFace and resized to 112x112.
This was… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/ms1mv3-wds-gz.MANUS-HaGRID
MANUS-HaGRID: HaGRID-derived Multimodal Annotated Naturalistic Hand Understanding Dataset
MANUS-HaGRID is the HaGRID/HaGRIDv2-derived subset of the Multimodal Annotated Naturalistic Hand Understanding (MANUS) dataset family. It provides multimodal annotations for naturalistic hand gesture understanding, including RGB images, hand crops, depth maps, 2D bounding boxes, estimated MANO-style hand mesh metadata, and multi-view mesh renderings where available.
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/QFun/MANUS-HaGRID.SPIDER-colorectal
SPIDER-COLORECTAL Dataset
SPIDER is a collection of supervised pathological datasets covering multiple organs, each with comprehensive class coverage. These datasets are professionally annotated by pathologists.
If you would like to support, sponsor, or obtain a commercial license for the SPIDER data and models, please contact us at models@hist.ai.
For a detailed description of SPIDER, methodology, and benchmark results, refer to our research paper:
📄 SPIDER: A Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/histai/SPIDER-colorectal.plate_effects
Multi-Source Domain Adaptation for Bioimaging Data (MSCDA-BioIm)
MSCDA-BioIm is a biomedical microscopy benchmark for evaluating test-time and in-context domain adaptation under realistic batch effects. Built from the large-scale JUMP-CP dataset, it targets mechanism-of-action (MoA) classification using five-channel images of compounds associated with eight well-defined MoA classes. The dataset is organized by experimental batches and imaging… See the full description on the dataset page: https://huggingface.co/datasets/anasanchezf/plate_effects.TreeOfLife-10M
Dataset Card for TreeOfLife-10M
Dataset Summary
With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/ScarlettChan/TreeOfLife-10M.imagenet-gen-sd1.5
ImageNet Generated using Stable Diffusion v1.5
The following repository mimics the size and class structure of the original ImageNet database. The classes can be found in the classes.txt file.
This dataset contains approximately 1300 images per class over 1000 classes for a total of 1.3 million images.
Here is an excerpt from classes.txt:
0 tench, Tinca tinca
1 goldfish, Carassius auratus
2 great white shark, white shark, man-eater, man-eating shark, Carcharodon caharias
3 tiger… See the full description on the dataset page: https://huggingface.co/datasets/ek826/imagenet-gen-sd1.5.NOAA_buoycams
NOAA Buoycam Dataset
This dataset provides NOAA buoycam images paired with meteorological observation data from the time the image was captured.
Dataset Details
Created using SeeSea (https://github.com/BrianOfrim/SeeSea)
Json Observation Data:
Key
Type
Unit
description
station_id
string
N/A
Id of the buoy
timestamp
string
YYYY_MM_dd_HHmm
Date/time of image & observation
description
string
N/A
Data
lat_deg
float
degrees [-90, 90]
Latitude… See the full description on the dataset page: https://huggingface.co/datasets/brianofrim/NOAA_buoycams.imagenet-w21-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with JPEG images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
Data Splits
The full ImageNet dataset has no defined splits. This release follows that and leaves everything in the train split.… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-wds.CDDB
Dataset Card for CDDB
Dataset Description
CDDB is a benchmark dataset introduced in the WACV 2023 paper A Continual Deepfake Detection Benchmark: Dataset, Methods, and Essentials.
It is designed for continual deepfake detection, where manipulated images from different deepfake generation sources arrive sequentially instead of being observed all at once.
The benchmark is intended to evaluate both:
binary deepfake detection (real vs. fake)
continual and incremental… See the full description on the dataset page: https://huggingface.co/datasets/nebula/CDDB.CUB_200_2011
Caltech-UCSD Birds-200-2011
Dataset Summary
This is a repackaged version of the Caltech-UCSD Birds-200-2011 dataset for convenient use with PyTorch's ImageFolder class.
Note: All credit goes to the original authors.
This upload only provides the same data in a different structure.
Website: https://www.vision.caltech.edu/datasets/cub_200_2011/
Paper: https://authors.library.caltech.edu/records/cvm3y-5hh21
200 categories dataset consists of 11,788 images (1.1GB).… See the full description on the dataset page: https://huggingface.co/datasets/birder-project/CUB_200_2011.cc0-textures
Dataset Card for CC0 Textures
Dataset Summary
This dataset contains 18,785 texture images from cc0-textures.com. It includes textures of wood, metal, concrete, fabric, stone, ceramic, and other materials. The original archives were downloaded, unpacked, and images were compressed using PNG optimization and JPEG quality compression (90%) to reduce file size while keeping good quality.
Languages
The dataset is monolingual:
English (en): Texture titles and tags… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cc0-textures.SPIDER-breast
SPIDER-BREAST Dataset
SPIDER is a collection of supervised pathological datasets covering multiple organs, each with comprehensive class coverage. These datasets are professionally annotated by pathologists.
If you would like to support, sponsor, or obtain a commercial license for the SPIDER data and models, please contact us at models@hist.ai.
For a detailed description of SPIDER, methodology, and benchmark results, refer to our research paper:
📄 SPIDER: A Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/histai/SPIDER-breast.
