Team Ai
Datasetpublic

OpenMOSS-Team/LearnFromMove

Learn from Move: LatentGUIWorld Benchmark LatentGUIWorld is the interactive GUI benchmark introduced in Learn from Move: the Next Step for GUI Agents. It contains 900 test episodes across six environments, with 150 episodes per environment and a 1280 × 720 viewport. Code and environment runtime Environments Configuration Task Episodes drag_egocentric Egocentric Drag 150 drag_exocentric Exocentric Drag 150 rotation_inner Inner Rotation 150… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/LearnFromMove.

sourceHugging Faceupdated 3d agoView on Hugging Face
1likes303downloads
Dataset Card

Learn from Move: LatentGUIWorld Benchmark

LatentGUIWorld is the interactive GUI benchmark introduced in Learn from Move: the Next Step for GUI Agents. It contains 900 test episodes across six environments, with 150 episodes per environment and a 1280 × 720 viewport.

Code and environment runtime

Environments

ConfigurationTaskEpisodes
drag_egocentricEgocentric Drag150
drag_exocentricExocentric Drag150
rotation_innerInner Rotation150
rotation_outerOuter Rotation150
ten_choice_egocentricEgocentric Ten-Choice150
ten_choice_exocentricExocentric Ten-Choice150

Drag places a colored shape into a matching outline. Ten-Choice identifies a target among ten candidates through hover-revealed text. Rotation uses a horizontal slider to align the inner or outer image region. Agents use interaction feedback to adapt their actions to each environment's dynamics.

The two environments in each task family share 150 paired scene identities, giving 450 scene pairs and 900 episodes. The pair_id identifies the shared scene; the episode_id identifies a particular environment episode.

All Hugging Face configurations expose a test split. The default all configuration combines the same six subsets into 900 rows. The distribution field identifies the 75 IID and 75 OOD episodes in each Drag environment. Rotation and Ten-Choice each use fixed test sets without an IID/OOD subdivision, so their distribution values are null.

The held-out factors follow the paper's test-set construction. Drag's IID episodes use training shape categories, while its OOD episodes use held-out shape categories. All Ten-Choice test scenes draw from 240 messages disjoint from the 80 training messages. Rotation uses 150 test background images disjoint from the 1,000 training backgrounds. Ten-Choice and Rotation therefore evaluate held-out content across their full test sets.

Load episode records

python
import json
from datasets import load_dataset

repo_id = "OpenMOSS-Team/LearnFromMove"
episodes = load_dataset(repo_id, "all", split="test")
drag = load_dataset(repo_id, "drag_egocentric", split="test")
episode = json.loads(drag[0]["episode_json"])
FieldMeaning
episode_id, pair_idEpisode and paired-scene identifiers
suite_id, variant, familyBenchmark suite, environment, and task family
instructionTask instruction
exploration_levelExploration category
distributioniid or ood for Drag; null for Rotation and Ten-Choice
canonical_case_pathScene metadata path relative to the dataset root
episode_jsonComplete runtime episode configuration serialized as JSON

The runtime consumes the full configuration, including hidden dynamics and success criteria. Agent observations consist of task instructions, screenshots, and the structured metadata selected by the runtime's observation contract.

Environment interaction

This dataset contains scene configurations, rendering assets, and the Ten-Choice HTML/JavaScript scenes. The complete environment runtime, mouse-action interface, and evaluation code are available in the GitHub repository.

Download and run the benchmark

python
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="OpenMOSS-Team/LearnFromMove",
    repo_type="dataset",
    local_dir="benchmark",
)

Install the latest environment runtime from the code repository, then pass the downloaded manifest to the environment or evaluation entry point:

bash
python -m six_environments list --manifest benchmark/manifest.json
python -m six_environments serve --manifest benchmark/manifest.json --port 8765
latentguiworld-eval --manifest benchmark/manifest.json \
  --model models/latentlearner --output results/latentlearner

Paths are relative to the directory where the commands are run. The evaluator scores the first release attempt and reports success rates per environment.

Files

text
README.md
LICENSE
NOTICE.md
manifest.json                 # Benchmark runtime manifest: 900 episodes
data/*.jsonl                  # Six Hugging Face subsets: 150 rows each
cases/drag/*/meta.json
cases/rotation/*/meta.json
cases/ten_choice/*/meta.json
cases/ten_choice/*/index.html
cases/ten_choice/*/assets/icon.svg
assets/rotation/*.jpg

The JSONL files provide a browsable view of the episodes in manifest.json. Scene paths in both representations resolve against the dataset root.

Attribution

The original release license is included in LICENSE. Rotation backgrounds originate from Open Images; their source references are retained in scene metadata, and image rights remain with their respective owners. See NOTICE.md.