datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FashionMnist_test
IllusionFashionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionFashionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each metadata row identifies a Fashion-MNIST target and can be paired across five image conditions: source-condition, illusion, filtered illusion, illusionless control, and filtered illusionless control.
The source-condition images originate from… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/FashionMnist_test.MNIST_test
IllusionMNIST — Test Set
Dataset summary
This repository contains the public test split of IllusionMNIST, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Every indexed example can be compared across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control images.
The source-condition images are sampled from MNIST and resized to 512 × 512 pixels. Illusion images were… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/MNIST_test.IllusionAnimals_test
IllusionAnimals — Test Set
Dataset summary
This repository contains the public test split of IllusionAnimals, introduced in Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. Each annotated example is paired across source-condition, illusion, filtered-illusion, illusionless-control, and filtered-illusionless-control conditions.
The animal source-condition images were generated with SDXL-Lightning. English scene descriptions and… See the full description on the dataset page: https://huggingface.co/datasets/VQA-Illusion/IllusionAnimals_test.OpenSDI_test
OpenSDI: Spotting Diffusion-Generated Images in the Open World
This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper:
Project Page: https://iamwangyabin.github.io/OpenSDI/
OpenSDID Dataset Highlights:
User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs.
Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.gui-odyssey-test
Dataset Card for GUI Odyssey (Test Split)
This is a FiftyOne dataset with 29426 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/gui-odyssey-test")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/gui-odyssey-test.isu-challenge-dataset
Dataset Card for ISU Challenge Dataset
Dataset Summary
ISU Challenge Dataset is a multi-modal in-cabin automotive dataset with a controlled synthetic core and paired real-reference scenes.
The synthetic core contains 1,000 synchronized Blender-rendered samples with:
RGB render
depth (EXR and PNG)
instance segmentation
canny edge map
structured scenario labels
An additional 60 paired real-reference scenes occupy sample_00000 through sample_00059. Each paired… See the full description on the dataset page: https://huggingface.co/datasets/ISU-Test/isu-challenge-dataset.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.mind2web_multimodal_test_domain
Dataset Card for "Cross-Domain" Test Split in Multimodal Mind2Web
Note: This dataset is the test split of the Cross-Domain dataset introduced in the paper.
This is a FiftyOne dataset with 4050 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_domain.mind2web_multimodal_test_task
Dataset Card for Multimodal Mind2Web "Cross-Task" Test Split
Note: This dataset is the test split of the Cross-Task dataset introduced in the paper.
This is a FiftyOne dataset with 1338 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_task.guiact_websingle_test
Dataset Card for GUIAct Web-Single Dataset - Test Set
This is a FiftyOne dataset with 1410 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/guiact_websingle_test")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/guiact_websingle_test.autotrain-data-ethnicity-test_v003
AutoTrain Dataset for project: ethnicity-test_v003
Dataset Description
This dataset has been automatically processed by AutoTrain for project ethnicity-test_v003.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<512x512 RGB PIL image>",
"target": 1
},
{
"image": "<512x512 RGB PIL image>",
"target": 3
}]… See the full description on the dataset page: https://huggingface.co/datasets/cledoux42/autotrain-data-ethnicity-test_v003.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.Mirage-Test
🌊 Mirage-Test Dataset
Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models.
It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics.
The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity.
📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.WebUOT-238-Test
Dataset Card for WebUOT-238-Test
This is a FiftyOne dataset with 238 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/WebUOT-238-Test")
# Launch the App
session = fo.launch_app(dataset)
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/WebUOT-238-Test.cpsc5800-hand-detection-test
Test Datasets and Model Weights for CPSC 5800 Final Project
Project repository: https://github.com/rohanphanse/CPSC5800-Final
We provide all test datasets created in Step 1 and weights for the YOLO and ResNet models trained during Steps 2-4 in our Hugging Face repository: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test
# Recommended: download dataset using git-xet (https://hf.co/docs/hub/git-xet)
brew install git-xet
git xet install
# Download datasets and… See the full description on the dataset page: https://huggingface.co/datasets/rohanphanse/cpsc5800-hand-detection-test.autotrain-data-testttt
AutoTrain Dataset for project: testttt
Dataset Description
This dataset has been automatically processed by AutoTrain for project testttt.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<113x220 RGB PIL image>",
"target": 2
},
{
"image": "<1280x720 RGB PIL image>",
"target": 2
}]
Dataset Fields
The… See the full description on the dataset page: https://huggingface.co/datasets/AdamOswald1/autotrain-data-testttt.vqa-test-studies
UCSF-PDGM Test Subset (VQA Demo)
A three-study subset of the UCSF Preoperative Diffuse Glioma MRI (UCSF-PDGM) dataset, mirrored here as test fixtures for a web-based visual question answering (VQA) demo. The full source dataset is publicly available on The Cancer Imaging Archive (TCIA).
Contents
Three preoperative brain MRI studies from patients with diffuse glioma:
UCSF-PDGM-0159_nifti/
UCSF-PDGM-0338_nifti/
UCSF-PDGM-0529_nifti/
Each study folder contains… See the full description on the dataset page: https://huggingface.co/datasets/chihhua0908/vqa-test-studies.testing_molmo2
Dataset Card for Voxel51/qualcomm-interactive-video-dataset
This is a FiftyOne dataset with 100 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/testing_molmo2")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_molmo2.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.testing-goldstandard-cuthill
Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset
Dataset Description
Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total).
There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed)
The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.autotrain-data-resnet50_test
AutoTrain Dataset for project: resnet50_test
Dataset Description
This dataset has been automatically processed by AutoTrain for project resnet50_test.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<1920x1920 RGB PIL image>",
"target": 2
},
{
"image": "<1080x721 RGB PIL image>",
"target": 2
}
]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/SRGui/autotrain-data-resnet50_test.NTIRE-RobustAIGenDetection-test-public
Test set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
📄 CVPR 2026 Workshop Paper · NTIRE 2026
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for trust… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-test-public.winml-test-set
WinML Test Set
Dataset Summary
WinML Test Set is an evaluation‑only collection for validating model accuracy and stability on Windows ML / DirectML / ONNX Runtime pipelines. It aggregates several permissively‑licensed sources and harmonizes schema for reproducible, regression‑grade testing across backends and versions. Not intended for training.
Intended Use
Accuracy and regression benchmarking of Windows ML / DirectML / ONNX Runtime pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/Futuremark/winml-test-set.testing_qwen3vl_on_video
Dataset Card for harpreetsahota/random_short_videos
This is a FiftyOne dataset with 412 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/testing_qwen3vl_on_video")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_qwen3vl_on_video.bdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.frakturline-testset
Fraktur/Other Text-Line — Test Set
A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman).
Developed for the Impresso digital humanities project.
Dataset Details
Property
Value
Task
Binary image classification
Classes
fraktur, other
Images per class
1 000
Total images
2 000
Image format
WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.photo-test
Anime vs. Live Action Film Image Classification
OmarK211/photo-test
A binary image classification dataset designed to distinguish between Anime (Label 0) and Live Action (Label 1) film frames/imagery. Images are prepared as standardized square RGB files with multi-pass synthetic training variants.
Source and task
The images were manually imported from the web into my local computer. Then manually uploaded to Google Colab
Preparation source: 24-679 Image Data… See the full description on the dataset page: https://huggingface.co/datasets/OmarK211/photo-test.testset
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.testing_deepseek_ocr
Dataset Card for Voxel51/document-haystack-10pages
This is a FiftyOne dataset with 250 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/testing_deepseek_ocr")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/testing_deepseek_ocr.UTM_Testing_Dataset
Dataset Description
This dataset was created to support Article VisionGauge: a computer vision model to detect and read U-tube manometers.
It consists of images of U-tube manometers constructed using a transparent PVC water level hose (5/16" × 1 mm) and flexible measuring tape with a length of 150 cm (60 inches). The manometric fluids represented in the dataset include water, oil, and dyed water. The dataset is intended for testing ML models for reading liquid column… See the full description on the dataset page: https://huggingface.co/datasets/claytonsds/UTM_Testing_Dataset.
