datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aimultiple-decision-models-browser
AIM Decision Model Browser Benchmark: 10-task sample
This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares 19 decision models (also called System One models) with 2 general LLMs on browser tasks. The decision models are Jev 1.13, Kev-4B, Kev-9B, Kev-27B, AutoJev-27B, Eikos-4B, Eikos-27B, decider-4b, decider-12b, GLiDE, Clef, Clef-Flash, Laya typed-decisions, GLiNER2.5-Decide, CLM-v0.1-8B, Solar Decide, Nimble 9B, Tev1 4B and Tev1… See the full description on the dataset page: https://huggingface.co/datasets/AIMultiple/aimultiple-decision-models-browser.nnetnav-exploration-browseruse
NNetNav exploration data → executed-action decisions with browser-use-style observations
Web-navigation decisions ("which action next?") derived from the raw exploration logs of
NNetNav (Murty, Zhu, Bahdanau, Manning: NNetNav: Unsupervised Learning of Browser
Agents Through Environment Interaction in the Wild), published as
ProKil/nnetnav-exploration-data (CC-BY-4.0,
revision ccb40bbe32). All credit for the explorations, the retroactive instructions and the
recordings goes to… See the full description on the dataset page: https://huggingface.co/datasets/johannhartmann/nnetnav-exploration-browseruse.BrowserBenchSee: https://github.com/Halluminate/browserbench
whisper-browser-benchmarks
whisper-browser-benchmarks
Measurements from a Whisper transcription pipeline running entirely inside a
browser tab: which audio and video containers the browser will actually decode,
how accurate the smallest usable Whisper size is on clean synthetic speech, how
long transcription takes relative to the length of the clip, what the first
load pulls over the wire, and what happens to clips longer than the model's
30-second window.
Everything here was measured, not quoted from a… See the full description on the dataset page: https://huggingface.co/datasets/ruanjiange/whisper-browser-benchmarks.building-permits-browser-data2d-webmcp-browser-focus
2D WebMCP Browser Focus (Prerelease)
What this is
This is an early test of whether agents need useful tool results to complete an accessible browser task.
The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer.
What happened
We ran each version 20 times with gpt-5-mini using low reasoning effort.
Tool result
Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.BrowserBenchSee: https://github.com/Halluminate/browserbench
