datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.GUIOdyssey
Dataset Card for GUIOdyssey
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Paper: https://arxiv.org/pdf/2406.08451
News⭐️
Latest version of GUIOdyssey released!🎉
This updated version features a larger dataset with 8,334 episodes, as well as richer semantic annotations. Compared to the previous version, we have added more fine-grained low-level instructions, image descriptions, action intentions, and context review for each step. Additionally, we provide bounding… See the full description on the dataset page: https://huggingface.co/datasets/hflqf88888/GUIOdyssey.go-swe-bench-v0
go_swe_bench v0 — real Go bug fixes, verified by the Go toolchain
246 tasks from 79 real Go repositories. Each task is a bug-fix commit whose co-committed test is red on the
parent and green on the fix. No LLM anywhere in the build.
Mined on 2026-09-19 from the GuildLM Go mining pipeline by inverting the filter that had thrown the tests
away (the pipeline was built for SFT data; a benchmark needs the opposite). Every task was verified twice
with go test: green at the commit (≥ 1… See the full description on the dataset page: https://huggingface.co/datasets/guildlm/go-swe-bench-v0.guildnouketsukejoudesugazangyouwaiyananodebosswosolotoubatsushiyoutoomoimasu
Bangumi Image Base of Guild No Uketsukejou Desu Ga, Zangyou Wa Iya Nanode Boss Wo Solo Toubatsu Shiyou To Omoimasu
This is the image base of bangumi Guild no Uketsukejou desu ga, Zangyou wa Iya nanode Boss wo Solo Toubatsu Shiyou to Omoimasu, we detected 64 characters, 4480 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/guildnouketsukejoudesugazangyouwaiyananodebosswosolotoubatsushiyoutoomoimasu.guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online medical sources. This dataset serves as a crucial component of the original training corpus of the Meditron Large Language Model (LLM). We publicly release a subset of 37K articles… See the full description on the dataset page: https://huggingface.co/datasets/epfl-llm/guidelines.Gui-agent
Gui-Agent — GUI trajectories in LIBERO/VLA format
Human GUI demonstrations from four sources, unified into a single VLA-style
intermediate representation and written as LIBERO-layout HDF5, so LIBERO/VLA
dataloaders run against GUI data unchanged.
raw source ──[adapter]──> GuiEpisode ──[writer]──> LIBERO-style HDF5
per-source the IR format- what you train on
only specific
25,872 episodes / 453,264 steps / 235 GB… See the full description on the dataset page: https://huggingface.co/datasets/Yushi123/Gui-agent.ATO-Australian-Tax-Rulings-and-Guidance
ATO Rulings & Guidance — Australian Tax Law, Structured for AI
67,000+ Australian Taxation Office documents as RAG-ready NDJSON/CSV — Edited Private Advice, public rulings and determinations, ATO Interpretative Decisions, practical compliance guidelines, taxpayer alerts, decision impact statements, practice statements and legislative instruments. Every document parsed into structured, typed fields for legal RAG, LLM fine-tuning, and tax research automation.
Machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/simplelex/ATO-Australian-Tax-Rulings-and-Guidance.GUI-Net-1M
Check more details at how to use this dataset at our repo
GUI-Net-1M is the dataset we keep running the pipeline introduced from TongUI paper.
Due to large file size, we have to split image files into parts. To do the extraction of images, please use the following script:
#!/bin/bash
# Directory containing the split files
SPLIT_DIR="/mnt/bofeidisk2/tmp/baidu_experience_full/images/split_parts_baidu_experience"
OUTPUT_DIR="merged_files"
# Create output directory if it doesn't… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-1M.easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPGUI_BASED_PLATFORMGUI-Net-Video
GUI-Net-Video
A collection of video URLs for GUI-related tasks, organized by platform or application.
This dataset provides the video URLs collected for TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents — a framework that turns multimodal web tutorials (including YouTube videos) into large-scale GUI agent trajectories. Each URL is paired with a GUI task description (the .txt filename) and grouped by platform/app.
Paper: arXiv:2504.12679… See the full description on the dataset page: https://huggingface.co/datasets/Bofeee5675/GUI-Net-Video.MolmoPoint-GUISyn
MolmoPoint-GUISyn
MolmoPoint-GUISyn is a large-scale synthetic dataset of 36K GUI screenshots with dense pointing annotations for training GUI grounding agents. Each screenshot is a realistic simulation of a digital environment (desktop apps, mobile apps, websites) generated entirely from code, with an average of 54 annotated UI elements per image.
The data is generated using the MolmoPoint-GUISyn pipeline, with Claude Sonnet 4.6 as the coding LLM.
Quick links:
Model:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoPoint-GUISyn.DATA_SOURCEguitarset
GuitarSet
GuitarSet v1.1.0 (Xi et al. 2018) on HuggingFace. 360 rows × 4 audio captures per row + canonical labels derived from the original JAMS. Three documented upstream errata are corrected (see below); the original file-based distribution lives on Zenodo with a permanent DOI. CC-BY 4.0.
Schema
Column
Type
Notes
track_id
string
e.g. 00_BN1-129-Eb_comp
player
int32
0–5
style
string
comp | solo
tempo_bpm
float64
key, mode, key_mode
string
e.g. Eb… See the full description on the dataset page: https://huggingface.co/datasets/jhartquist/guitarset.GUIDE-dataset
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
GUI Unbiasing via Instructional-video Driven Expertise
Accepted to ECCV 2026
Project Page |
Paper |
GitHub
This dataset supports the accepted ECCV 2026 paper "GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation".
Overview
GUIDE (GUI Unbiasing via Instructional-Video… See the full description on the dataset page: https://huggingface.co/datasets/sharryXR/GUIDE-dataset.gui_actor_hfaf-esc-guidelines-v1
Dataset Card for "af-esc-guidelines-v1"
More Information needed
guiltycrown
Bangumi Image Base of Guilty Crown
This is the image base of bangumi Guilty Crown, we detected 30 characters, 2278 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/guiltycrown.GenOmni3D
3D GLB Mesh Dataset
This dataset contains 3D mesh objects in GLB format.
GUI-AIMA-multiturnGUIrilla-Task
GUIrilla-Task
Ground-truth Click & Type actions for macOS screenshots
Dataset Summary
GUIrilla-Task pairs real macOS screenshots with free-form natural-language instructions and precise GUI actions.
Every sample asks an agent either to:
Click a specific on-screen element, or
Type a given text into an input field.
Targets are labelled with bounding-box geometry, enabling exact evaluation of visual-language grounding models.
Data were gathered automatically by… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/GUIrilla-Task.guide_and_masterguiltygearstrivedualrulers
Bangumi Image Base of Guilty Gear Strive: Dual Rulers
This is the image base of bangumi Guilty Gear Strive: Dual Rulers, we detected 37 characters, 1811 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/guiltygearstrivedualrulers.gui-odyssey-1kGUIDEMobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT
This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text.
Repository Layout
.
├── GUI-agent-main/ # Data annotation scripts and examples
├── eval/ # Evaluation assets
│ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.GUIOdyssey
cua-lite/GUIOdyssey
cua-lite preprocessed version of GUIOdyssey (hflqf88888/GUIOdyssey). A long-horizon cross-app Android mobile dataset of 8,334 task trajectories over ~128k screenshots. Produces two cohorts: use (multi-step agent episodes) and understanding (per-step screen captioning from the source description annotations).
Origin
https://huggingface.co/datasets/hflqf88888/GUIOdyssey
Load via datasets
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/cua-lite/GUIOdyssey.guiowl-aw-mix-full
GUI-Owl AndroidWorld SFT Mix — FULL / generalization
Purpose: AndroidWorld (116-task) SFT for a GUI-Owl-1.5-2B block-diffusion VLA.
This dataset is an action-balanced, source-mixed SFT corpus assembled from five Android
GUI-agent trajectory sources. It is built for in-domain supervised fine-tuning ahead of RL.
The mix deliberately includes AndroidWorld task-family coverage (via the openmobile anchor,
whose app field holds AW task-family names) and accepts in-domain overlap by… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/guiowl-aw-mix-full.schema_guided_dialogThe Schema-Guided Dialogue (SGD) dataset contains 18K multi-domain task-oriented
dialogues between a human and a virtual assistant, which covers 17 domains
ranging from banks and events to media, calendar, travel, and weather. The
language presents in the datset is only English. The SGD dataset provides a
challenging testbed for a number of tasks in task-oriented dialogue, including
language understanding, slot filling, dialogue state tracking and response
generation. For the creation of the SGD dataset, they developed a multi-domain
dialogue simulator that generates dialogue outlines over an arbitrary combination
of APIs, dialogue states and system actions. Then, they used a crowd-sourcing
procedure to paraphrase these outlines to natural language utterances. This novel
crowd-sourcing procedure preserves all annotations obtained from the simulator and
does not require any extra annotations after dialogue collection.
