Team Ai
Datasetpublic

AIMultiple/aimultiple-decision-models-browser

AIM Decision Model Browser Benchmark: 10-task sample This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares 19 decision models (also called System One models) with 2 general LLMs on browser tasks. The decision models are Jev 1.13, Kev-4B, Kev-9B, Kev-27B, AutoJev-27B, Eikos-4B, Eikos-27B, decider-4b, decider-12b, GLiDE, Clef, Clef-Flash, Laya typed-decisions, GLiNER2.5-Decide, CLM-v0.1-8B, Solar Decide, Nimble 9B, Tev1 4B and Tev1… See the full description on the dataset page: https://huggingface.co/datasets/AIMultiple/aimultiple-decision-models-browser.

sourceHugging Facecc-by-nc-nd-4.0updated 5d agoView on Hugging Face
1likes220downloads
Dataset Card

AIM Decision Model Browser Benchmark: 10-task sample

This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares 19 decision models (also called System One models) with 2 general LLMs on browser tasks. The decision models are Jev 1.13, Kev-4B, Kev-9B, Kev-27B, AutoJev-27B, Eikos-4B, Eikos-27B, decider-4b, decider-12b, GLiDE, Clef, Clef-Flash, Laya typed-decisions, GLiNER2.5-Decide, CLM-v0.1-8B, Solar Decide, Nimble 9B, Tev1 4B and Tev1 0.8B. The LLMs are GPT-6 Astra and Gemini 3.8 Flash. Kev-4B ran twice, on a rented GPU and through OpenRouter, so the results cover 22 runs. The article does not rank Solar Decide, Nimble 9B, the two Tev1 models and Laya on browser tasks, because their services accept at most 26 options per question or limit input length; their results are included here. The other 40 tasks are withheld so they can be reused for later runs.

GPT-6 Astra, Gemini 3.8 Flash, Jev, Kev-9B and Laya ran on 22 September 2026, with the browser on a Mac and Kev-9B on a rented A100 GPU. GLiNER and CLM ran on 25 September on rented RTX 4090 GPUs. Kev-4B ran on 29 September on a rented RTX A6000 GPU and through OpenRouter. On 30 September, Solar Decide ran through OpenRouter, and Nimble 9B and both Tev1 models ran through Ollama 0.35.0 on a rented RTX A6000 GPU. GLiDE, Clef and Clef-Flash ran on 2 October through their providers' APIs from the Mac. Kev-27B, AutoJev-27B, Eikos-27B, Eikos-4B, decider-4b and decider-12b ran on 2 and 3 October on rented A100 GPUs. The Jev rows are from the 22 September run.

Article: https://aimultiple.com/decision-models

Files

  • —tasks.jsonl has one task per line: start URL, instruction, accepted final URLs and pass conditions.
  • —results.jsonl has one row per task and run: pass or fail, how the attempt ended, seconds, number of browser actions and run date.
  • —media/ has two demo clips of the models working the same task at real speed.

How tasks are scored

An attempt passes only if the model declares the task done within the limits and an independent check of the final page confirms every condition. The URL must be one of the accepted final URLs, every listed selector condition must hold, and the page must be fully loaded. A model's own claim of success does not count.

Conditions use CSS selectors. text is the element's whitespace-normalized text, texts is the list over all matching elements, value is a form value, count is the number of matches and visible means rendered and not hidden by CSS.

Limits per attempt were 900 seconds, 60 browser actions and 120 decision calls. Every attempt started in a fresh browser session and ran once.

Sample

IDSiteType
B04Books to ScrapeCategory switch to product page
W01WikipediaArticle lookup
Q06Quotes to ScrapeDependent filters
Q08Quotes to ScrapeInfinite scroll
V07Web ScraperNative dropdown
H05Scrape This SiteAJAX table
C06ScrapingCourseProduct options
M02ScrapeMeSort
D04web-scraping.devLoad more
T02TestPagesForm submit

The tasks were chosen to cover every site in the benchmark and a mix of outcomes. H05 was passed in 13 of the 22 runs and T02 in 3. Pass rates in this sample do not match the full benchmark: Jev, Kev-9B and Solar Decide pass 5 of these 10 tasks but completed 17, 20 and 10 of all 50.

Runtime

All models used the same open-source runtime, browser-use/jev-ultrafast. At each step it turns the page into text plus a numbered list of visible controls, and the model picks an operation and a target. Values typed into form fields come from a separate text model, Mercury 2.5, which was available to every participant. Clef, Clef-Flash, decider and Eikos require at least two options per question, so a question with one option was sent with that option twice and the answer was mapped back to it.

Practice sites change over time. A task that fails today because a site changed is a site problem, not a model result.

Demo clips

Five models on W01 (top row Jev, Kev-9B, Laya; bottom row Gemini 3.8 Flash, GPT-6 Astra and a results card):

[image]

Four models on Q06, filtering quotes by author and tag:

[image]

The clips come from separate demo runs, not from the scored attempts in results.jsonl.