datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.rlbenchfail_test_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.rlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.rlbenchfail_val_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_val_dataset.NautData
NautData
Paper | Project Page | Code
NautData is a large-scale underwater instruction-following dataset containing 1.45 million image-text pairs. It was constructed to bridge the gap in large-scale underwater multi-task instruction-tuning datasets, which are crucial for advancing underwater scene understanding methods. The dataset enables the development and thorough evaluation of underwater Large Multimodal Models (LMMs).
This dataset was introduced in the paper NAUTILUS: A Large… See the full description on the dataset page: https://huggingface.co/datasets/H-EmbodVis/NautData.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting instances, which follows the COCO format… See the full description on the dataset page: https://huggingface.co/datasets/eelianafang/MOUNT-Cattle.polyvore-outfits
Polyvore Outfits (Refactored Version)
This repository provides a refactored version of the Polyvore Outfits dataset, originally introduced in the paper "Learning Type-Aware Embeddings for Fashion Compatibility" by Mariya I. Vasileva et al.
📌 Overview
The goal of this refactoring is to improve usability and developer experience. While the core data remains identical to the original, the file structure and JSON schemas have been standardized to make it easier to load and… See the full description on the dataset page: https://huggingface.co/datasets/owj0421/polyvore-outfits.geoguesser-tasks
GeoGuesser Task Splits
Task indexes for the GeoGuesser OpenEnv environment.
Each line is one episode: an ordered list of panorama frames with coordinates,
headings and capture dates, plus the sequence and contributor it came from.
Split
Tasks
Countries
Frames
Fully mirrored
eval
200
73
4673
200/200
train
3452
130
80179
3448/3452
What a task is
These files carry metadata only, not imagery. Every frame's coordinates,
heading and capture date are… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/geoguesser-tasks.GeoChrono-Data
ChronoBench & ChronoInstruct
ChronoBench is a comprehensive, multi-dimensional, and multi-granularity benchmark for high-resolution long-temporal remote sensing understanding. It decomposes long-term remote sensing understanding into a four-level cognitive hierarchy — from Land Cover Perception through Temporal Recognition and Long-Term Memory to Spatio-Temporal Reasoning — comprising 12 sub-tasks and 17,689 rigorously validated QA pairs derived from 3,469 high-resolution… See the full description on the dataset page: https://huggingface.co/datasets/Davidup1/GeoChrono-Data.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/vpasx/lgg-mri-segmentation-research.ur5fail_test_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.ur5fail_train_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting… See the full description on the dataset page: https://huggingface.co/datasets/y1665065879/MOUNT-Cattle.ur5fail_val_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_val_dataset.vernier
vernier
Error bars on a dataset vendor's quality claim. Build AI publishes hand-visibility and
active-manipulation rates for Egocentric-10K / Egocentric-100K, judged once by
gemini-2.5-flash with no human gold, no interval, and no test that the judge scores a factory
floor and a home kitchen on the same scale. This release is the data behind an independent,
pre-registered measurement of that claim: human labels against a written rubric, a live
open-weights judge on the same… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/vernier.VisionEncoder-Eval-ReproDataA Strong Baseline for Evaluating Vision Encodersin Multimodal Large Language Models
Yilin Yang1,* ·
Jun-Tao Tang2,* ·
Kengyi Wang3 ·
Siyuan Su3 ·
Gaoyong Luo4 ·
Mingda Chen1,†
1School of Artificial Intelligence, Shanghai Jiao Tong University
2Nanjing University ·
3Fudan University ·
4Independent Researcher
*Equal contribution. †Corresponding author.… See the full description on the dataset page: https://huggingface.co/datasets/336labs/VisionEncoder-Eval-ReproData.bdv2fail_train_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.AgroTools
Anonymous question-4 release
This repository contains an anonymized release folder for the question-4 split.
Included files
metadata.jsonl: normalized table for the dataset viewer
images/: image assets referenced by the dataset
assets/AppleSizeEstimate/: depth .npy assets referenced by selected samples
question-4.original.json: original source question file
question_taxonomy_summary.md: taxonomy and template notes
Split
test: 539 samples
Schema… See the full description on the dataset page: https://huggingface.co/datasets/AgroTools/AgroTools.pad-auto-solver-reviewed
PAD Reviewed Dataset
Canonical reviewed PAD board/orb artifacts for dw-indie/pad-auto-solver-reviewed. This repository
contains immutable reviewed package revisions and does not contain raw captures,
training runs, checkpoints, or model binaries.
Packages exported: 28
Active catalog datasets: 14
Catalog schema: 3
Layout
packages/<dataset_id>.tar: deterministic self-contained reviewed package
catalog.json: active revision heads and coverage summary… See the full description on the dataset page: https://huggingface.co/datasets/dw-indie/pad-auto-solver-reviewed.bdv2fail_val_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_val_dataset.bdv2fail_test_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.OpenHotels
OpenHotels
OpenHotels is a large-scale hotel image retrieval benchmark built from hotel-room imagery and associated hotel metadata. The dataset is designed for hotel-scale retrieval: given a query image, a system must retrieve the matching hotel from a large gallery containing both true matching classes and many distractor hotel classes.
Dataset Structure
The release contains tar-sharded image files under shards/ and four metadata files:
shards/… See the full description on the dataset page: https://huggingface.co/datasets/imagingforgood/OpenHotels.radread-public-results
RadRead — public results
Rollout-level results for RadRead, a benchmark of frontier models reading 150
radiographs. Every row is one graded model read: 5 saved rollouts per study
per model, scored by a deterministic grader (no judge model).
A read passes only when every required checklist finding, lesion box (the grader's
IoU / centre / containment test), lexical diagnosis check and action-set membership
check match the reference rubric. No partial credit inside a study;… See the full description on the dataset page: https://huggingface.co/datasets/tirandazdylan/radread-public-results.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.CPRT-Bench
Dataset Card for CPRT-Bench
CPRT-Bench is a benchmark dataset for assessing privacy risk in images, designed to model privacy as a graded and composition-dependent phenomenon.
Dataset Details
Dataset Description
The dataset contains approximately 6.7K images annotated with:
Ordinal severity levels (4 levels of privacy risk)
Continuous risk scores (fine-grained privacy assessment)
All images are sourced from the VISPR (Visual Privacy Advisor). CPRT-Bench… See the full description on the dataset page: https://huggingface.co/datasets/timtsapras23/CPRT-Bench.Light-RAG-Marketing-Assets-Agent
🖼️ Light RAG Marketing Assets Agent — Pre-ingested Data
Pre-ingested LightRAG knowledge graph and vector data from 420 marketing images
analyzed with Gemini Vision API (gemini-3.5-flash) and processed through GPT-4o
for entity extraction and relationship mapping.
GitHub repo: 0xrphl/Light-RAG-Marketing-Assets-Agent
📊 Dataset Statistics
Metric
Value
Source images
420 (JPG/PNG/WebP)
Text chunks
2,095 (5 per image: core, visual, people/setting… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/Light-RAG-Marketing-Assets-Agent.MONITRS
MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
Dataset Description
Paper: NeurIPS 2025 (Spotlight)
Contact: revankar@cs.cornell.edu
MONITRS contains ~10,000 FEMA disaster events with temporal Sentinel-2 satellite imagery, natural language captions from news articles, geotagged locations, and question-answer pairs for disaster monitoring research.
Supported Tasks
Event classification
Temporal grounding
Location grounding
Visual… See the full description on the dataset page: https://huggingface.co/datasets/ShreelekhaR/MONITRS.MOUNT-Cattle
Updates/News 📣
🎉 News (Feb. 2026): The dataset paper FSMC-Pose has been accepted for CVPR 2026 Findings!
🔗 News: Please find the open-source dataset on Hugging Face: MOUNT-Cattle.
🔥 Downloads reached 2.4k within 7 days of release.
📌 Overview
Mounting posture is an important visual indicator of estrus in dairy cattle. MOUNT-Cattle is a mounting dataset, covering 1,176 mounting… See the full description on the dataset page: https://huggingface.co/datasets/chenziyue-cattle/MOUNT-Cattle.
