datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-research-berkeley-webagent
Berkeley WebAgent Experiment Artifacts
Native GEPA, CLUE and ACE experiment logs and available actor trajectory evidence.
Files require manual access approval. Request access with your Hugging Face account.
Results and complete evidence snapshot — September 23, 2026
Combined experiment summary: GEPA, CLUE and ACE; method × site for WebArena.
Non-WebArena repeat scores and variance, including completed ACE ≤50k-token evaluations.
WebArena / GoBrowse progress and… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-webagent.seeclick-web-commercial-mlx
SeeClick Web Commercial Dataset (MLX-VLM Format)
Commercial-use friendly GUI grounding dataset from SeeClick Web data.
Apache 2.0 licensed - safe for commercial applications.
Dataset Description
This dataset contains ~20k examples for training Vision-Language Models to predict
click coordinates given a screenshot and instruction. Derived from SeeClick Web
crawled data (Apache 2.0).
Key Features
License: Apache 2.0 (commercial use allowed)
Format: MLX-VLM… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/seeclick-web-commercial-mlx.pi-webPi sessions of https://github.com/woxQAQ/pi-web exported by pi-share-hf
mimo-v2.6-distill-qwen9b-webdev-chunk00
MiMo V2.6 Distill Qwen 9B Webdev — Chunk 00
Public intermediate artifact from the MiMo V2.6 teacher-capture series:
9 complete, filter-selected synthetic web-development examples (prompt +
full teacher HTML completion + grader score) for SFT and knowledge
distillation into small code-capable language models.
Teacher: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Prompts attempted: 10
Complete filtered examples: 9
Generation: Kaggle 2x T4, FP16, one candidate, max_new_tokens=12288
Local… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/mimo-v2.6-distill-qwen9b-webdev-chunk00.WebUI-COCO-876
WebUI-COCO-876: Annotated Webpage Screenshots for UI Region & Landmark Detection
Short description
876 desktop webpage screenshots annotated with COCO bounding boxes over ARIA-style landmark regions and common UI components (cookie dialogs, popovers, captcha, buttons, etc.)
Dataset details
Modality: image (webpage screenshots)
Annotation format: COCO-style JSON (bounding boxes + class IDs, optional semantic attributes)
Number of images: 876
Label… See the full description on the dataset page: https://huggingface.co/datasets/jileklu/WebUI-COCO-876.web-fetch-harness-traces
Native web fetch harness traces
Separate, lightly sanitized native JSONL traces comparing URL-fetch behavior in Claude Code and Codex CLI against:
https://huggingface.co/datasets/nyu-mll/glue
Captured on 2026-09-15. No shell HTTP client, browser automation, or MCP fetcher was used.
Files
data/claude-code.jsonl: Claude Code's stream-json events.
data/codex.jsonl: Codex's persisted native rollout JSONL.
Both traces are exposed together in the default subset and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/web-fetch-harness-traces.rephrased-web-data-quality-study
Rephrased Web Data Quality Study
LLM-as-judge evaluation of ~4,000 examples from HuggingFaceFW/finephrase (1,000 sampled per split, 86 dropped due to judge parse failures, 3,914 successfully evaluated).
Judge: Claude Sonnet 4.6 via OpenRouter | Cost: ~$45
Quality Scores (1-5 scale)
Metric
FAQ (n=965)
Table (n=979)
Tutorial (n=976)
Math (n=994)
Faithfulness
1.82
1.72
1.90
1.49
Info preservation
1.93
1.64
1.99
1.47
Appropriateness
3.54
2.87
2.48
1.67… See the full description on the dataset page: https://huggingface.co/datasets/ratishsp/rephrased-web-data-quality-study.matrix_operations
░▒▓██████████████▓▒░ ░▒▓██████▓▒░▒▓████████▓▒░▒▓███████▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░▒▓██████████████▓▒░ ░▒▓██████▓▒░
░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░ ░▒▓█▓▒░ ░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░
░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░ ░▒▓█▓▒░ ░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░
░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓████████▓▒░ ░▒▓█▓▒░ ░▒▓███████▓▒░░▒▓█▓▒░░▒▓██████▓▒░░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓████████▓▒░
░▒▓█▓▒░░▒▓█▓▒░░▒▓█▓▒░▒▓█▓▒░░▒▓█▓▒░… See the full description on the dataset page: https://huggingface.co/datasets/webxos/matrix_operations.webcode2m-improved-promptsBCI-grid-movement-v1
BCI Grid Movement Intent Dataset
UNDER DEVELOPMENT for TESTING purposes
This dataset was generated with the gym located in /gym/ folder of this repo.The dataset contains synchronized movement intent data collected from a grid training environment designed for simulated
BCI research. Dataset for multi-label classification of WASD movement intents from 12 simulated EEG channels in grid environment.
This BCI Intent Data Study (conceptual early design) is for training machine… See the full description on the dataset page: https://huggingface.co/datasets/webxos/BCI-grid-movement-v1.WEBPRMBENCH
WebPRMBench
The first comprehensive evaluation benchmark for Web Process Reward Models
Published at ICLR 2026
Paper | Code | Website | Collection | Demo
Overview
WebPRMBench is the first comprehensive evaluation benchmark dedicated to Web Process Reward Models (WebPRMs). It evaluates how well a reward model can judge the quality of web agent actions during long-horizon web navigation. Each instance presents a web state (page context, trajectory history, user… See the full description on the dataset page: https://huggingface.co/datasets/ZYao720/WEBPRMBENCH.agent_trajectory_reviewsionicocean
___ ___ ___ ___ ___ ___ ___ ___
___ / /\ / /\ ___ / /\ / /\ / /\ / /\ / /\ / /\
/__/\ / /::\ / /::| /__/\ / /::\ / /::\ / /::\ / /::\ / /::\ / /::|
\__\:\ / /:/\:\ / /:|:| \__\:\ / /:/\:\ /… See the full description on the dataset page: https://huggingface.co/datasets/webxos/ionicocean.megamath-web-pro-max-splittedsd-webui-forge-configwebgpt_comparisons_ko
original dataset: openai/webgpt_comparisons
wavebender_dataset
_ _ __ _ _ ____ ____ ____ _ _ ____ ____ ____
( \/\/ ) /__\( \/ )( ___)( _ \( ___)( \( )( _ \( ___)( _ \
) ( /(__)\\ / )__) ) _ < )__) ) ( )(_) ))__) ) /
(__/\__)(__)(__)\/ (____)(____/(____)(_)\_)(____/(____)(_)\_)
OVERVIEW
UNDER DEVELOPMENT
This dataset was generated using the WAVEBENDER app by webXOS, located in the /generator/ folder of this repo. Download WAVE BENDER
to create your own similar datasets.… See the full description on the dataset page: https://huggingface.co/datasets/webxos/wavebender_dataset.arachnid_RL
___ ___ ___ ___ ___ ___
/\ \ /\ \ /\ \ /\__\ /\ \ /\ \ _____
/::\ \ /::\ \ /::\ \ /:/ / \:\ \ \:\ \ ___ /::\ \
/:/\:\ \ /:/\:\__\ /:/\:\ \ /:/ / \:\ \ \:\ \ /\__\ /:/\:\ \
/:/ /::\ \ /:/ /:/ / /:/ /::\ \ /:/ /… See the full description on the dataset page: https://huggingface.co/datasets/webxos/arachnid_RL.webcode2m-scored-prompts-gpt-osssnipercell_RL
_____ _ _ ___________ ___________ _____ _____ _ _
/ ___| \ | |_ _| ___ \ ___| ___ \ / __ \| ___| | | |
\ `--.| \| | | | | |_/ / |__ | |_/ / | / \/| |__ | | | |
`--. \ . ` | | | | __/| __|| / | | | __|| | | |
/\__/ / |\ |_| |_| | | |___| |\ \ | \__/\| |___| |____| |____
\____/\_| \_/\___/\_| \____/\_| \_| \____/\____/\_____/\_____/
SNIPER CELL -… See the full description on the dataset page: https://huggingface.co/datasets/webxos/snipercell_RL.drone_fsd_dataset
Drone FSD Dataset
A Single Sample training run of: 1 epoch, 4 iterations, 198 steps in drone navigation in a 60×60 room with 15 static + 12 floating obstacles.
This dataset was generated with the MIRROR IDE by webXOS. Download the app in the /mirror/ folder to train your own similar datasets.
Final performance (after 2456 frames):
- Best time: 43.821 s
- Success rate: 0.0% (reached SE corner in best run but did not complete full pattern)
- Collisions: 0 in final… See the full description on the dataset page: https://huggingface.co/datasets/webxos/drone_fsd_dataset.mimo-v2.6-distill-qwen9b-webdev-16
MiMo V2.6 Distill Qwen 9B WebDev Teacher Responses
Small, quality-filtered web-development SFT dataset generated from
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B.
This is an independently generated, public research dataset for testing
knowledge distillation, supervised fine-tuning, and webdev model training.
It is not an official Xiaomi dataset.
What is included?
14 complete HTML teacher responses
16 original prompts sampled from the MiMo-V2.6 webdev training split… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/mimo-v2.6-distill-qwen9b-webdev-16.webXOS_magnet_dataset
___ ___ ___ ___ ___
/\ \ /\__\ _____ /| | /\ \ /\__\
_\:\ \ /:/ _/_ /::\ \ |:| | /::\ \ /:/ _/_
/\ \:\ \ /:/ /\__\ /:/\:\ \ |:| | /:/\:\ \ /:/ /\ \
_\:\ \:\ \ /:/ /:/ _/_ /:/ /::\__\ __|:|__| /:/ \:\ \ /:/ /::\ \
/\ \:\ \:\__\ /:/_/:/ /\__\ /:/_/:/\:|__| /::::\__\_____… See the full description on the dataset page: https://huggingface.co/datasets/webxos/webXOS_magnet_dataset.open-web-vectors-manifest
Open Web Vector Initiative — Site Manifest
Per-site metadata for every site in the Open Web Vector Initiative, including
what each site told us about AI use on the day we asked.
The initiative — how the permission gate works, and what we will and will
not publish: https://divinci.ai/open-web-vectors/
The live directory — search the corpus, chat with any site in it, or claim
your own: https://divinci.ai/www-rag/
This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.memgym-rm-scenario-ood-webarena
MemGym-RM Scenario-OOD — WebArena V2
Description
MemGym-RM-Scenario-OOD-WebArena is a held-out evaluation set designed to test whether MemRM generalizes to a completely different agent domain (WebArena browser tasks) not seen during training (which used SWE-Gym software engineering tasks). Each row is a trajectory step from a WebArena agent running under one of three memory cohorts, with a scenario_swap perturbation applied.
OOD axis: Agent scenario / task domain (SWE-Gym… See the full description on the dataset page: https://huggingface.co/datasets/MemGym/memgym-rm-scenario-ood-webarena.LeroyDyer___Spydaz_Web_AI_AGI_R1_Math_AdvancedStudent-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_AGI_R1_Math_AdvancedStudent
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_AGI_R1_Math_AdvancedStudent
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_AGI_R1_Math_AdvancedStudent-details.mimo-v2.6-distill-qwen9b-webdev-full
MiMo V2.6 Distill Qwen 9B Webdev — Full Curated Dataset
Consolidated, strictly filtered teacher-distillation dataset from Chunks 00–12.
Teacher: XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Prompts attempted: 130
Accepted teacher samples: 108
Strict SFT samples: 107
Filter: complete <!DOCTYPE html> document ending in </html> plus local webdev grading
Generation: Kaggle 2x T4, FP16, one candidate, max_new_tokens=12288
One malformed sample from Chunk 04 is intentionally excluded. Each… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/mimo-v2.6-distill-qwen9b-webdev-full.LeroyDyer___Spydaz_Web_AI_ChatQA_003-details
Dataset Card for Evaluation run of LeroyDyer/_Spydaz_Web_AI_ChatQA_003
Dataset automatically created during the evaluation run of model LeroyDyer/_Spydaz_Web_AI_ChatQA_003
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer___Spydaz_Web_AI_ChatQA_003-details.tip-of-my-tongue-known-item-search-triplets
The TOMT-KIS-TRIPLETS (tip-of-my-tongue-known-item-search triplets) Dataset
TOMT-KIS-TRIPLETS is a refined subset of the broader TOMT-KIS dataset, curated for advanced applications. Given that Wikipedia and IMDb are besides YouTube the most commonly appearing domains in links on the answer path (also see in TOMT-KIS the column links_on_answer_path), we focus specifically on these sources for higher relevancy and usability. By leveraging a Wikipedia dump and SPARQL Wikipedia Query… See the full description on the dataset page: https://huggingface.co/datasets/webis/tip-of-my-tongue-known-item-search-triplets.web-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.
