datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FashionMnist_test
IllusionFashionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionFashionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row identifies a Fashion-MNIST target and can be paired across five image conditions: source-condition, illusion, filtered illusion, illusionless control, and filtered illusionless control.
The source-condition images originate from… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_test.MNIST_test
IllusionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images.
The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.IllusionAnimals_test
IllusionAnimals — Test Set
Dataset summary
This repository contains the public test split of IllusionAnimals, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each annotated example is paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions.
The animal source-condition images were generated with SDXL-Lightning. English scene descriptions and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_test.OpenSDI_test
OpenSDI: Spotting Diffusion-Generated Images in the Open World
This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper:
Project Page: https://iamwangyabin.github.io/OpenSDI/
OpenSDID Dataset Highlights:
User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs.
Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.gui-odyssey-test
Dataset Card for GUI Odyssey (Test Split)
This is a FiftyOne dataset with 29426 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/gui-odyssey-test")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/gui-odyssey-test.isu-challenge-dataset
Dataset Card for ISU Challenge Dataset
Dataset Summary
ISU Challenge Dataset is a multi-modal in-cabin automotive dataset with a controlled synthetic core and paired real-reference scenes.
The synthetic core contains 1,000 synchronized Blender-rendered samples with:
RGB render
depth (EXR and PNG)
instance segmentation
canny edge map
structured scenario labels
An additional 60 paired real-reference scenes occupy sample_00000 through sample_00059. Each paired… See the full description on the dataset page: https://huggingface.co/datasets/ISU-Test/isu-challenge-dataset.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.mind2web_multimodal_test_domain
Dataset Card for "Cross-Domain" Test Split in Multimodal Mind2Web
Note: This dataset is the test split of the Cross-Domain dataset introduced in the paper.
This is a FiftyOne dataset with 4050 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_domain.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.guiact_websingle_test
Dataset Card for GUIAct Web-Single Dataset - Test Set
This is a FiftyOne dataset with 1410 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/guiact_websingle_test")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/guiact_websingle_test.Mirage-Test
🌊 Mirage-Test Dataset
Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models.
It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics.
The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity.
📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.cpsc5800-hand-detection-test
Test Datasets and Model Weights for CPSC 5800 Final Project
Project repository: https://github.com/rohanphanse/CPSC5800-Final
We provide all test datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test
# Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet)
brew install git-xet
git xet install
# Download datasets and… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test.testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.NTIRE-RobustAIGenDetection-test-public
Test set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
📄 CVPR 2026 Workshop Paper · NTIRE 2026
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for trust… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-test-public.winml-test-set
WinML Test Set
Dataset Summary
WinML Test Set is an evaluation‑only collection for validating model accuracy and stability on Windows ML / DirectML / ONNX Runtime pipelines. It aggregates several permissively‑licensed sources and harmonizes schema for reproducible, regression‑grade testing across backends and versions. Not intended for training.
Intended Use
Accuracy and regression benchmarking of Windows ML / DirectML / ONNX Runtime pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/Futuremark/winml-test-set.frakturline-testset
Fraktur/Other Text-Line — Test Set
A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman).
Developed for the Impresso digital humanities project.
Dataset Details
Property
Value
Task
Binary image classification
Classes
fraktur, other
Images per class
1 000
Total images
2 000
Image format
WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.photo-test
Anime vs. Live Action Film Image Classification
OmarK211/photo-test
A binary image classification dataset designed to distinguish between Anime (Label 0) and Live Action (Label 1) film frames/imagery. Images are prepared as standardized square RGB files with multi-pass synthetic training variants.
Source and task
The images were manually imported from the web into my local computer. Then manually uploaded to Google Colab
Preparation source: 24-679 Image Data… See the full description on the dataset page: https://huggingface.co/datasets/OmarK211/photo-test.testset
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.testing_deepseek_ocr
Dataset Card for Voxel51/document-haystack-10pages
This is a FiftyOne dataset with 250 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/testing_deepseek_ocr")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_deepseek_ocr.UTM_Testing_Dataset
Dataset Description
This dataset was created to support Article VisionGauge: a computer vision model to detect and read U-tube manometers.
It consists of images of U-tube manometers constructed using a transparent PVC water level hose (5/16" × 1 mm) and flexible measuring tape with a length of 150 cm (60 inches). The manometric fluids represented in the dataset include water, oil, and dyed water. The dataset is intended for testing ML models for reading liquid column… See the full description on the dataset page: https://huggingface.co/datasets/claytonsds/UTM_Testing_Dataset.test
Tiny overlapping ImageFolder demo
A tiny synthetic repository for testing Hugging Face Dataset Viewer.
12 lossless WebP images, each exactly 128×128.
full: 8 train, 2 validation, 2 test.
core: 4 train, 1 validation, 1 test.
core is a subset of full.
Images exist once in images/.
manifest.csv is the canonical manifest.
splits/ contains lightweight source indices.
Root full_*.csv and core_*.csv are materialized metadata files used by ImageFolder.
Expected Viewer… See the full description on the dataset page: https://huggingface.co/datasets/Krows7/test.autotrain-data-test_row2
AutoTrain Dataset for project: test_row2
Dataset Description
This dataset has been automatically processed by AutoTrain for project test_row2.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<316x316 RGB PIL image>",
"target": 1
},
{
"image": "<316x316 RGB PIL image>",
"target": 3
}]
Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/Efimov6886/autotrain-data-test_row2.GenAI-RealEstate-TestSet
🏙️ GenAI Real Estate Test Set (Track B)
Dataset for the MenaML Winter School 2026 Challenge.
📊 Dataset Structure
This dataset contains 1,000 images split evenly between:
Authentic: Real estate photography from the Places365 dataset.
Manipulated: Synthetically generated deepfake artifacts (Inpainting, Diffusion Noise, GAN Grids).
🕵️ How to Use
This dataset is designed for testing forensic detection models.
test_nested_dataset
Image Classification - Monochrome Or Not
2 labels, 260 samples in total, listed as the following:
Label
Samples
Sample #0
Sample #1
Sample #2
Sample #3
Sample #4
Sample #5
Sample #6
Sample #7
monochrome
4 (1.5%)
N/A
N/A
N/A
N/A
colored
256 (98.5%)
diagram-eval-access-test
Diagram Evaluation Access Test
This public one-image dataset tests the external-access workflow planned for the
Diagram Evaluation project. It is not a research dataset release.
Dataset structure
The repository uses Hugging Face's ImageFolder layout:
data/
train/
metadata.csv
sample-diagram.png
The metadata records the image's provenance, license, and intended use. A production
release can use the same contract with sharded Parquet or WebDataset files.… See the full description on the dataset page: https://huggingface.co/datasets/abhisheklalwani96/diagram-eval-access-test.TesticulUS
TesticulUS
A multicenter testicular ultrasound resource for segmentation and
parenchymal-inhomogeneity classification research.
Collection
Annotations
Availability
Two Italian clinical centers
Segmentation mask for every image
Via Ditto only
First 860 images
Parenchymal-inhomogeneity classification label
Via Ditto only
Overview
TesticulUS combines testicular ultrasound images acquired at two Italian
clinical centers. Every image is paired with… See the full description on the dataset page: https://huggingface.co/datasets/AImageLab-Zip/TesticulUS.HowToEat-test
HowToEat: Hand-Object Interaction and Eating Action in Eating Scenarios
HowToEat is an image dataset for analysing eating behaviour. It provides:
Hand-object interaction + eating face detection (hand_object_detection): 95,190 images with 190,333 hand instances (box, left/right side, contact state, and the box and category of the held object) and 151,620 face instances (box, eating / not eating).
Eating action recognition (eating_recognition): 6,280 manually labelled faces… See the full description on the dataset page: https://huggingface.co/datasets/thxplz/HowToEat-test.astrobridge-yse-test-dataset-v2
AstroBridge YSE external test dataset v2
This dataset contains 266 object-disjoint, spectroscopically labeled YSE DR1 transients. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15.
V2 shortens the YSE forced-photometry time coverage to resemble the alert-photometry coverage of the AstroBridge BTS training dataset. For each object, it retains the smallest inclusive time interval that contains every positive measurement with flux/uncertainty at least 5 and at least five… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset-v2.imagenette_segmented_testai-detector-benchmark-test-data
🎯 AI Detector Benchmark Test Dataset
A comprehensive benchmark dataset for testing AI image detection models.
📊 Dataset Summary
Total Images: 700
AI-Generated: 250 images (from 5 different generators)
Real Images: 450 images (from 9 diverse datasets)
Perfect for:
✅ Testing AI detection models
✅ Creating leaderboards
✅ Comparing model performance
✅ Benchmarking new approaches
🤖 AI Generators Included
Generator
Images
Accuracy Baseline
FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.
