datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.GUI-CC
GUI-CC
GUI-CC is a benchmark for evaluating the contextual consistency of GUI world models when
they are used as agent environments rather than as isolated next-screen predictors.
A GUI world model predicts the next interface given the current screenshot and an action.
When that prediction is fed back as the next state, the rollout must stay coherent: app identity,
navigation history, created entities, selected options, and action affordances all have to remain
mutually… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/GUI-CC.omniact-gui-trajectories
OmniACT
OmniACT is a GUI trajectory dataset with single-step action traces grounded in screenshots.
Dataset Structure
.
├── README.md
├── .gitattributes
├── data/
│ └── train.jsonl
├── observations/
│ └── OmniACT_pilot_*/000/screenshot.jpg
└── env_meta/
└── OmniACT_pilot_*/000/metadata.json
Each row in data/train.jsonl is one trajectory. The main image path is stored in the top-level image field, and the same relative path is also used inside… See the full description on the dataset page: https://huggingface.co/datasets/Dhscl/omniact-gui-trajectories.gui-primitives
GUI-Primitives
A controlled minimal-pair diagnostic benchmark for the elementary spatial primitives
that GUI click instructions depend on.
Author(s): Md Abrar Jahin, Md Rizwan Parvez
Accepted at EMNLP 2026 (Main Conference).
VLM-based computer-use agents fail largely at grounding — turning a language
instruction into a click coordinate. Existing spatial-reasoning benchmarks
(What's-Up, BLINK, CV-Bench, VSR) use natural photographs, not screenshots, and none
isolate which… See the full description on the dataset page: https://huggingface.co/datasets/kagnlp/gui-primitives.GUI-Critic-TrainGUI-C2-4K
