datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UI-Genie-Agent-16kThis repository contains the Trajectory dataset from the paper UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based
Mobile GUI Agents.
Github: https://github.com/Euphoria16/UI-Genie
UIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/LiamLian0727/UIIS.Uni-GUI-Desktop-1
Uni-GUI-Desktop-1
A large-scale desktop GUI agent trajectory dataset, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,685
Total Steps
~36K
Platform
Desktop (1920x1080)
Applications
10 categories
Coordinate System
Normalized to [0, 999]
Applications
App
Description
chrome
Web browsing tasks
gimp… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Desktop-1.ui-vision
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Introduction
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.Uni-GUI-Mobile
Uni-GUI-Mobile
A mobile GUI agent trajectory dataset, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
871
Total Steps
~14K
Platform
Mobile (1080x2400)
Applications
10 categories
Coordinate System
Normalized to [0, 999]
Applications
App
Description
calendar
Calendar event management
chrome
Mobile… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Mobile.UIEBUni-GUI-OpenMobile
Uni-GUI-OpenMobile
A mobile GUI agent trajectory dataset collected on open-source Android applications via AndroidWorld, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,640
Total Steps
~25.9K
Platform
Android Mobile (1080x2400)
Applications
19 open-source apps
Coordinate System
Normalized to [0, 1000]… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenMobile.MedVidBench
MedVidBench: A Benchmark for Medical Video Understanding
Introduced in the paper: MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding (CVPR 2026).
📄 Paper: arxiv.org/abs/2512.06581
🌐 Project Page: uii-ai.github.io/MedGRPO
💻 Code: UII-AI/MedGRPO-Code
🤗 Model: UII-AI/uAI-NEXUS-MedVLM-1.0a-7B-RL
🎮 Demo: UII-AI/MedGRPO-Demo
📊 Leaderboard: UII-AI/MedVidBench-Leaderboard
Dataset Description
MedVidBench is a test benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/UII-AI/MedVidBench.wave-uiLICENSE
Uni-GUI-Desktop-2
Uni-GUI-Desktop-2
A desktop GUI agent trajectory dataset collected on OSWorld environments, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
1,247
Total Steps
~14.8K
Platform
Desktop (1920x1080)
Applications
10 categories
Coordinate System
Normalized to [0, 1000]
Applications
App
Count
Description… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Desktop-2.UI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.wave-ui-25k
WaveUI-25k
This dataset contains 25k examples of labeled UI elements. It is a subset of a collection of ~80k preprocessed examples assembled from the following sources:
WebUI
RoboFlow
GroundUI-18K
These datasets were preprocessed to have matching schemas and to filter out unwanted examples, such as duplicated, overlapping and low-quality datapoints. We also filtered out many text elements which were not in the main scope of this work.
The WaveUI-25k dataset includes the original… See the full description on the dataset page: https://huggingface.co/datasets/agentsea/wave-ui-25k.Uni-GUI-OpenCUA
Uni-GUI-OpenCUA
A post-processed desktop GUI agent trajectory dataset derived from OpenCUA, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Overview
Uni-GUI-OpenCUA contains 832 trajectories with ~14K interaction steps across 11 desktop applications and task categories. The raw OpenCUA trajectories have been cleaned, filtered, and post-processed through our Unified Cross-Platform Data… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenCUA.MobileWorld-Eval-Results
MobileWorld-Eval-Results
Evaluation results of the UI-MOPD trained model (Qwen3-VL-8B-Thinking) on a mobile agent benchmark. Contains full execution trajectories including screenshots, marked action visualizations, and task outcomes across 117 diverse mobile tasks.
Evaluation Summary
Metric
Value
Model
Qwen3-VL-8B-Thinking
Total Tasks
117
Successful
12
Success Rate
10.3%
Action Space
mobile_use
Max Steps
50
Avg Steps
32.5
Total Steps
~3.8K… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/MobileWorld-Eval-Results.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.ui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/ijlewis/ui-navigation-corpus.UIEB
Dataset Card for "UIEB"
More Information needed
easyr1-103k-4MP-jedi-ui-vision-gta1-data
easyr1-103k-4MP-jedi-ui-vision-gta1-data
Merged dataset composed of the following sources:
datasets/easyr1-63k-nores-jedi-fix-synced-ui-vision-manually-labeled-icon-data-from-yt-4MP (63031 samples in split train)
datasets/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b (39943 samples in split train)
Summary
Generated on: 2025-09-18 06:29:16 UTC
Split: train
Column strategy: intersection
Samples after merge: 102974
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-103k-4MP-jedi-ui-vision-gta1-data.AndroidControl-Star
AndroidControl*
A curated step-level evaluation subset extracted from AndroidControl, used for static mobile GUI understanding evaluation in UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Overview
AndroidControl* contains 4,260 step-level records from 781 Android trajectories. Each record includes the trajectory identifier, step index, high-level task goal, per-step instruction, normalized action, screenshot path, screenshot… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/AndroidControl-Star.UIIS-semanticUIISThis dataset is proposed by the ICCV 2023 paper "WaterMask: Instance Segmentation for Underwater Imagery", specific parameters about the dataset can be viewed in the paper
The Underwater Image Instance Segmentation (UIIS) dataset contains 4,628 images with pixel-level annotations in seven categories used for the underwater instance segmentation task. The dataset is organized in MS COCO format and the annotation files and images for training and testing are in UDW files.
Updatae:… See the full description on the dataset page: https://huggingface.co/datasets/CoffeeZongzi/UIIS.awg_retina_tomography_uiui-navigation-corpus
User Interface (Navigation) Corpus
Overview
This dataset serves as a collection of various images of, videos and metadata of mobile (both iOS and Android) and web user interfaces as well as tags and text extractions associated to them.
Dataset also includes user interface navigation annotations and videos related to them. One of the possible use cases of this dataset is training a UI navigation agent.
Dataset Structure
The resources of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/teleren/ui-navigation-corpus.mobile-ui-design
Dataset: Mobile UI Design Detection
Introduction
This dataset is designed for object detection tasks with a focus on detecting elements in mobile UI designs. The targeted objects include text, images, and groups. The dataset contains images and object detection boxes, including class labels and location information.
Dataset Content
Load the dataset and take a look at an example:
>>> from datasets import load_dataset
>>>> ds =… See the full description on the dataset page: https://huggingface.co/datasets/mrtoy/mobile-ui-design.OSWorld-Eval-Results
OSWorld-Eval-Results
Evaluation results of the UI-MOPD trained model (Qwen3-VL-8B-Thinking) on the OSWorld benchmark. Contains full execution trajectories including screenshots, action logs, and task outcomes.
Evaluation Summary
Metric
Value
Model
Qwen3-VL-8B-Thinking
Total Tasks
359
Successful
126
Success Rate
35.1%
Action Space
pyautogui
Observation
Screenshot (1920x1080)
Max Steps
50
Coordinate
Relative
Per-Application… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/OSWorld-Eval-Results.UiPad
UiPad - UI Parsing and Accessibility Dataset
📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned.
Curated by: MacPaw Way Ltd.
Language(s): Mostly EN, UA
License: MIT
Overview
UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.ui-vision-grounding-4MPOpenViVQA-dataset
OpenViVQA: Open-domain Vietnamese Visual Question Answering
The OpenViVQA dataset contains 11,000+ images with 37,000+ question-answer pairs which introduces the Text-based Open-ended Visual Question Answering in Vietnamese. This dataset is publicly available to the research community in the VLSP 2023 - ViVRC shared task challenge. You can access the dataset as well as submit your results to evaluate on the private test set on the Codalab evaluation system.
Link to the OpenViVQA… See the full description on the dataset page: https://huggingface.co/datasets/uitnlp/OpenViVQA-dataset.uirp_drill_yolo_target_v2_20260826This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/hugging-son/uirp_drill_yolo_target_v2_20260826.gate-gsplat-assets
gate_scenes_export
Self-contained bundle of two sagesplat scenes (left/right) plus the object
extractions in their shared (joint) mocap frame.
Layout
gate_scenes_export/
├── README.md
├── MANIFEST.txt
├── left_scene/
│ ├── mocap_processed/ # data dir (referenced by config.yml as "./mocap_processed")
│ │ ├── images/ # 242 corrected RGB frames
│ │ ├── transforms.json # mocap-frame… See the full description on the dataset page: https://huggingface.co/datasets/uild42/gate-gsplat-assets.
