datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TreeOfLife-200M
Dataset Card for TreeOfLife-200M
If you are looking for the original release TreeOfLife-200M dataset, as used in training BioCLIP 2 and presented the paper, please see Revision a8f38b4. The dataset, as presented here, was used to train BioCLIP 2.5 Huge; it completes the dataset cleaning process and resolves an issue where Observation.org occurrences were not included in the training data.
With 233 million images representing 933,798 taxa across the tree of life, TreeOfLife-200M… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-200M.OpenFake
Dataset Card for OpenFake
OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.GroMo25
GroMo25: Multiview Time-Series Plant Image Dataset for Age Estimation and Leaf Counting
Dataset Summary
GroMo25 is a multiview, time-series plant image dataset designed for plant age estimation (in days) and leaf counting tasks in precision agriculture. It contains high-quality images of four crop species — Wheat, Okra, Radish, and Mustard — captured over multiple days under controlled conditions. Each plant is photographed from 24 angles across 5 vertical levels per day… See the full description on the dataset page: https://huggingface.co/datasets/MrigLabIITRopar/GroMo25.oxford-iiit-pet
The Oxford-IIIT Pet Dataset
Description
A 37 category pet dataset with roughly 200 images for each class. The images have a large variations in scale, pose and lighting.
This instance of the dataset uses standard label ordering and includes the standard train/test splits. Trimaps and bbox are not included, but there is an image_id field that can be used to reference those annotations from official metadata.
Website: https://www.robots.ox.ac.uk/~vgg/data/pets/… See the full description on the dataset page: https://huggingface.co/datasets/timm/oxford-iiit-pet.animal-clef-2026
AnimalCLEF26 Kaggle Competition Dataset
This is a HuggingFace mirror of the official AnimalCLEF26 competition dataset. Images have been repackaged into one zipfile per split, which include additional metadata that makes the dataset easier to use with HuggingFace. Otherwise, no files have been changed.
Loading
from datasets import load_dataset
dataset = load_dataset("BVRA/animal-clef-2026")
print(dataset["train"][0]["image"])
Documentation
For… See the full description on the dataset page: https://huggingface.co/datasets/BVRA/animal-clef-2026.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.imagenet1k-256-wdsThis is imagenet1k in webdataset format. Images are stored as jpg files. Every image has been resized to a maximum side length of 256. That means that if an image in the original dataset was 1000 by 500, the new size will be 256 by 128. Images with a maximum side length of under 256 were not resized.
The total size of all dataset files is 57.8 GB, there are 1,281,167 rows in the training split and 50,000 rows in the validation split.
zendo-synthetic-data
Zendo Synthetic Visual Reasoning Dataset
Synthetic Zendo-style scenes with associated rules and per-scene tensor
representations. Each scene either follows ("positive", label=1) or violates
("negative", label=0) a rule that is given in natural language and as a Prolog
query.
Splits
split
scenes
train
56475
test
3344
rules total
3439
Layout
images/<split>/<batch>/<rule_id>/<scene_id>.png — rendered scene… See the full description on the dataset page: https://huggingface.co/datasets/sophia1ch/zendo-synthetic-data.MIC21
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/18029/
MIC21
Original description
One of the processing tasks for large multimodal data streams is automatic image description (image classification, object segmentation and classification). Although the number and the diversity of image datasets is constantly expanding, still there is a huge demand for more datasets in terms of variety of domains and object classes covered. The goal of the… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/MIC21.lfw
LFW HF-ready
This folder packages the local LFW (Labeled Faces in the Wild) images as a
Hugging Face imagefolder dataset with the canonical 10-fold verification
pairs file.
Layout
lfw/
├── README.md
├── pairs.csv
└── train/
├── images/<shard>/<file>.jpg
└── metadata.csv
metadata.csv columns
file_name: relative image path used by ImageFolder, e.g. images/000/Aaron_Eckhart_0001.jpg.
label: numeric identity label.
label_name / identity: identity name.… See the full description on the dataset page: https://huggingface.co/datasets/marcelohaps/lfw.svg-benchmark
Rapidata Static SVG Generation Benchmark
Built by Rapidata.
This dataset contains 1,918,367 human responses, collected with the
Rapidata Python SDK, comparing how well 42 frontier LLMs generate
static SVGs from text prompts. Each row is a head-to-head comparison between two models' renders of
the same prompt, scored by human annotators on one of three questions (Preference, Coherence, Alignment).
The SVGs are produced as raw <svg> markup by the models, rasterized to 768×768 PNGs… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/svg-benchmark.Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.fish-vista
Dataset Card for Fish-Visual Trait Analysis (Fish-Vista)
Note that the '</Use this dataset>' option will only load the CSV files. To download the entire dataset, including all processed images and segmentation annotations, refer to Instructions for downloading dataset and images.
See Example Code to Use the Segmentation Dataset
Figure 1. A schematic representation of the different tasks in Fish-Vista Dataset.
Instructions for downloading dataset… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fish-vista.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.cola
COLA: Compose Objects Localized with Attributes
Self-contained Hugging Face port of the COLA benchmark from the paper
"How to adapt vision-language models to Compose Objects Localized with Attributes?".
📄 Paper: https://arxiv.org/abs/2305.03689
🌐 Project page: https://cs-people.bu.edu/array/research/cola/
💻 Original code & data: https://github.com/ArijitRay1993/COLA
This repository bundles the benchmark annotations as Parquet files and the referenced
images as regular files… See the full description on the dataset page: https://huggingface.co/datasets/array/cola.beans
Dataset Card for Beans
Dataset Summary
Beans leaf dataset with images of diseased and health leaves.
Supported Tasks and Leaderboards
image-classification: Based on a leaf image, the goal of this task is to predict the disease type (Angular Leaf Spot and Bean Rust), if any.
Languages
English
Dataset Structure
Data Instances
A sample from the training set is provided below:
{
'image_file_path':… See the full description on the dataset page: https://huggingface.co/datasets/AI-Lab-Makerere/beans.svgrepo
Dataset Card for SVGRepo Icons
Dataset Summary
This dataset contains a large collection of Scalable Vector Graphics (SVG) icons sourced from SVGRepo.com. The icons cover a wide range of categories and styles, suitable for user interfaces, web development, presentations, and potentially for training vector graphics or icon classification models. Each icon is provided under a specific open-source or permissive license, clearly indicated in its metadata. The SVG… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/svgrepo.InpaintCOCO
InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects)
Dataset Summary
A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object.
Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance.
Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.TAIX-Ray
TAIX-Ray Dataset
TAIX-Ray is a comprehensive dataset of approximately 200k bedside chest radiographs from around 50k intensive care patients at University Hospital Aachen, Germany, collected between 2010 and 2024.
Trained radiologists provided structured reports at the time of acquisition, assessing key findings such as cardiomegaly, pulmonary congestion, pleural effusion, pulmonary opacities, and atelectasis on an ordinal scale.
Code & Details
The code for data… See the full description on the dataset page: https://huggingface.co/datasets/TLAIM/TAIX-Ray.PlantVillage
PlantVillage Dataset
The PlantVillage Dataset is an open access repository of 54,306 images of healthy and diseased plant leaves, collected to advance research in automated plant disease diagnosis. It covers 14 crop species and 26 diseases.
This dataset was introduced in the paper "Using Deep Learning for Image-Based Plant Disease Detection" by Mohanty et al. (2016).
Quick Start
The dataset comes with pre-defined 80/20 train/test splits that preserve the leaf grouping… See the full description on the dataset page: https://huggingface.co/datasets/mohanty/PlantVillage.efficientnet-v2-l-adv-dataset
Perturb Adversarial Images
Verified adversarial examples for efficientnet_v2_l (torchvision/EfficientNet_V2_L_Weights.IMAGENET1K_V1), produced by the
Perturb network. Each row is one clean image together with all of its
verified adversarial versions: images that are imperceptibly different from the original
(L∞ ≤ 0.03 in [0,1] pixel scale) yet change the model's top-1 prediction.
This dataset grows continuously. New rows are appended as the network produces them and uploaded in… See the full description on the dataset page: https://huggingface.co/datasets/perturb-ai/efficientnet-v2-l-adv-dataset.webvid-10MeurosatRedistributed without modification from https://github.com/phelber/EuroSAT.
EuroSAT100 is a subset of EuroSATallBands containing only 100 images. It is intended for tutorials and demonstrations, not for benchmarking.
DeepTumorVQA_2.0
DeepTumorVQA v2
3D abdominal-CT diagnostic Visual Question Answering benchmark with 42
clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K
training pool). Includes pre-extracted 2D and video modalities, 20K agent
training trajectories with tool-use traces, and a paper-locked leaderboard.
Resources
📄 Paper (arXiv)
https://arxiv.org/abs/2605.09679
💻 Code (GitHub)
https://github.com/Schuture/DeepTumorVQA
🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.nih-chest-xray-14
Description
The NIH ChestX-ray14 dataset (NIH Clinical Center), an extension of ChestX-ray8 from the CVPR 2017 paper. It has 112,120 frontal-view chest X-rays of 30,805 unique patients. Each image is labelled with any of 14 thoracic findings, text-mined from the associated radiology reports with NLP.
'
ChestX-ray dataset comprises 112,120 frontal-view X-ray images of 30,805 unique patients with the text-mined fourteen disease image labels (where each image can have multi-labels)… See the full description on the dataset page: https://huggingface.co/datasets/timm/nih-chest-xray-14.FakeCOCO
FakeCOCO dataset
Using 10 SOTA text-to-image models to generate fake images based on COCO captions
over 1M images
These models include:
SD15
SD21
SDXL
SD3
Playground2.5
PixArt alpha
PixArt sigma
unidiffuser
Flux.1
Stable Cascade
IndoLepAtlas
IndoLepAtlas — Indian Lepidoptera & Host Plants Dataset
A large-scale computer vision dataset of Indian butterflies, moths, and their larval host plants. Sourced from ifoundbutterflies.org with public CC-licensed photographs.
Inspired by: iNaturalist | Domain: Indian Wildlife & Biodiversity
Dataset Overview
Butterflies
Host Plants
Total
Species
961
127
1,088
Images
60,641
703
61,344
Source
ifoundbutterflies.org
ifoundbutterflies.org
—… See the full description on the dataset page: https://huggingface.co/datasets/Butterfree/IndoLepAtlas.
