datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimultiple-decision-models-browser
AIM Decision Model Browser Benchmark: 10-task sample
This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares 19 decision models (also called System One models) with 2 general LLMs on browser tasks. The decision models are Jev 1.13, Kev-4B, Kev-9B, Kev-27B, AutoJev-27B, Eikos-4B, Eikos-27B, decider-4b, decider-12b, GLiDE, Clef, Clef-Flash, Laya typed-decisions, GLiNER2.5-Decide, CLM-v0.1-8B, Solar Decide, Nimble 9B, Tev1 4B and Tev1… See the full description on the dataset page: https://huggingface.co/datasets/AIMultiple/aimultiple-decision-models-browser.2d-webmcp-browser-focus
2D WebMCP Browser Focus (Prerelease)
What this is
This is an early test of whether agents need useful tool results to complete an accessible browser task.
The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer.
What happened
We ran each version 20 times with gpt-5-mini using low reasoning effort.
Tool result
Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.
