Team Ai
Datasetpublic

AmineHA/WebArena-Verified

WebArena-Verified Dataset description WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible evaluation of web agents across multiple realistic websites. Sources GitHub repository: webarena-verified Original WebArena benchmark: webarena.dev Splits full: 812 rows hard: 258 rows Tasks per site Counts below are task counts grouped by category. Tasks with more than one site are grouped… See the full description on the dataset page: https://huggingface.co/datasets/AmineHA/WebArena-Verified.

sourceHugging Faceupdated 8mo agoView on Hugging Face
2likes327downloads
Dataset Card

WebArena-Verified

Dataset description

WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible evaluation of web agents across multiple realistic websites.

Sources

Splits

  • —full: 812 rows
  • —hard: 258 rows

Tasks per site

Counts below are task counts grouped by category. Tasks with more than one site are grouped under multi-category to avoid double-counting.

Full split

SiteTasks
gitlab180
map109
reddit106
shopping_admin182
shopping187
wikipedia0
homepage0
multi-category48

Hard split

SiteTasks
gitlab57
map0
reddit42
shopping_admin55
shopping56
wikipedia0
homepage0
multi-category48

Schema notes

The table below reflects inferred column types from the full split during artifact generation.

ColumnType
sitesList(Value('string'))
task_idValue('int64')
intent_template_idValue('int64')
start_urlsList(Value('string'))
intentValue('string')
intent_templateValue('string')
instantiation_dictValue('string')
evalValue('string')
revisionValue('int64')

Metadata

  • —Version: v1.2.3
  • —Git commit: 6473f72db5dcefc97b5725b59e734504edc28a21
  • —Generated at (UTC): 2026-02-07T22:44:22Z
  • —Dataset hash: 0e90f3feabe68cb9c7285e6989e36862a47622e7ac81b562f3126152350cb5ac
  • —License: Apache-2.0
  • —Language: en
  • —Task categories: web-navigation, information-retrieval

Notes

  • —Expected split counts: full=812, hard=258
  • —If custom split names are not supported by the Dataset Viewer, use one config per split mapped to train as a compatibility fallback.