Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rmems /browser-tool-use-trajectories Browser Tool Use Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/browser-tool-use-trajectories.text1K<n<10K1 likes765 downloads18d agoHugging Face02AIMultiple /aimultiple-decision-models-browser AIM Decision Model Browser Benchmark: 10-task sample This is 10 of the 50 tasks from AIMultiple's decision model browser benchmark. The benchmark compares 19 decision models (also called System One models) with 2 general LLMs on browser tasks. The decision models are Jev 1.13, Kev-4B, Kev-9B, Kev-27B, AutoJev-27B, Eikos-4B, Eikos-27B, decider-4b, decider-12b, GLiDE, Clef, Clef-Flash, Laya typed-decisions, GLiNER2.5-Decide, CLM-v0.1-8B, Solar Decide, Nimble 9B, Tev1 4B and Tev1… See the full description on the dataset page: https://huggingface.co/datasets/AIMultiple/aimultiple-decision-models-browser.tabularothern<1K1 likes220 downloads5d agoHugging Face03WootzappLab /browser-agent-tasks Browser Agent Tasks Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence. Files tasks.jsonl: one task definition per line. Each task contains: task_id: stable identifier category: task family instruction: complete instruction given to the browser agent start_url: suggested public starting page stopping_condition: when the agent must stop constraints: safety and scope restrictions These tasks… See the full description on the dataset page: https://huggingface.co/datasets/WootzappLab/browser-agent-tasks.imagen<1K0 likes208 downloads10d agoHugging Face04TIGER-Lab /BrowserAgent-Data BrowserAgent ChatML Dataset (SFT/RFT) This dataset contains ChatML-style multi-turn dialogues for a browser agent task. The data is prepared as JSON Lines so it can be previewed directly with the Hugging Face Hub Data Visualizer and loaded with the datasets library. Links Paper Github Files sft.jsonl — SFT split (one JSON object per line) rft.jsonl — RFT split (one JSON object per line) Schema Each record is a JSON object containing: messages:… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/BrowserAgent-Data.texttext-generation10K<n<100K5 likes207 downloads11mo agoHugging Face05merve /browser-tasks Browser tasks Live-browser tasks both pseudogenerated from large models and also taken from only from the starting_url field of Halluminate/BrowserBench. Each source page was revisited live at 800×600. Up to three tasks were generated from the current page observation: information, navigation, and interaction. Historical prompts, results, screenshots, and ground-truth URLs were not used. The canonical dataset contains all 876 task slots from 292 source pages. A row is accepted… See the full description on the dataset page: https://huggingface.co/datasets/merve/browser-tasks.imagequestion-answeringn<1K0 likes98 downloads1d agoHugging Face06mutantweb /browser-agent-failure-corpus Browser Agent Failure Corpus — case-row view This view contains 30 controlled local regression cases from one owner-produced deterministic run on 2026-08-31 using Qwen/Qwen3-8B-GGUF:Q4_K_M, llama.cpp b10103-c588c4f47, the Codex in-app browser and the historical Mutant Web guard v2, at temperature 0, reasoning budget 0 and at most eight steps per case. It preserves model proposals, guard decisions, executed actions and recorded outcomes for inspectability and reuse. Allowed… See the full description on the dataset page: https://huggingface.co/datasets/mutantweb/browser-agent-failure-corpus.textn<1K1 likes78 downloads1mo agoHugging Face07cklxx /laya-browser-suite-c Suite C — 27 multi-step browser tasks on held-out real websites A small live benchmark for web agents: 27 tasks on 27 real sites, 19 of them needing 4 or more actions (search, then filters, sort, pagination, opening the right result). Every task has a programmatic success check on the final page (URL parameters, title, visible text or observed form state), so no LLM judge is needed. Built to evaluate cklxx/laya-browser: none of these domains appears in any of its training… See the full description on the dataset page: https://huggingface.co/datasets/cklxx/laya-browser-suite-c.textn<1K0 likes71 downloads11d agoHugging Face08ProCreations /betterwright-agentic-browser-50k BetterWright Agentic Browser — 6,093-row stopped checkpoint This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings. Contents 5,971 BetterWright demonstrations and 122 Playwright demonstrations. 32 task domains and 21 browser feature categories. Harness-shaped conversations… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/betterwright-agentic-browser-50k.texttext-generation1K<n<10K0 likes53 downloads1mo agoHugging Face09merve /browserbench-live BrowserBench Live Tasks This private dataset contains 292 live-browser task-generation records derived only from the starting_url field of Halluminate/BrowserBench at revision aa56ce5e6331425c29878037fa8c169507deecdc. Pages were visited live at generation time. Historical prompts, results, and ground-truth URLs were not used. Every source row is retained, including rejected pages, with its rejection reason. Fresh page screenshots are stored under screenshots/. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/merve/browserbench-live.imagen<1K0 likes50 downloads1mo agoHugging Face10DataCreatorAI /tool-calling-browser-agent-tasks Dataset Card Created by: DataCreator AI Overview Tool Calling for Agentic Tasks with Multi-Step Workflows contains 1,062 synthetic multi-turn conversations between a user and an AI assistant. The examples primarily focus on practical agentic tasks such as train ticket booking, dynamic form filling, and payment processing. It provides diverse scenarios including successful execution, context retrieval, tool integration, and failure recovery. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/tool-calling-browser-agent-tasks.text-generation1K<n<10K2 likes47 downloads7mo agoHugging Face11ishagarg1103 /browser-agent-tasks Browser Agent Tasks Moderate, multi-step browser tasks for collecting agent trajectories and evaluating screenshot, DOM, and DOM-diff evidence. Files tasks.jsonl: one task definition per line. Each task contains: task_id: stable identifier category: task family instruction: complete instruction given to the browser agent start_url: suggested public starting page stopping_condition: when the agent must stop constraints: safety and scope restrictions These tasks… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/browser-agent-tasks.textn<1K0 likes47 downloads1mo agoHugging Face12chrismat01 /BrowserSec-Bench-PreviewBrowserSec-Bench v0.1 — Free Preview Browser-validated web security benchmark for evaluating AI security systems. This repository contains a free 36-record preview of BrowserSec-Bench v0.1. The complete commercial edition contains 1,064 browser-executed and validated security experiments covering: CORS — 500 scenarios COOP / COEP / CORP — 264 scenarios Content Security Policy — 300 scenarios The benchmark is designed primarily for evaluating security-focused AI systems, security agents… See the full description on the dataset page: https://huggingface.co/datasets/chrismat01/BrowserSec-Bench-Preview.textn<1K0 likes39 downloads12d agoHugging Face13Vendex /agentic-tooluse-computer-browsertextn<1K0 likes38 downloads3mo agoHugging Face14parthkl /Agent-browser-tasktext1K<n<10K0 likes31 downloads4mo agoHugging Face15accesslint /2d-webmcp-browser-focus 2D WebMCP Browser Focus (Prerelease) What this is This is an early test of whether agents need useful tool results to complete an accessible browser task. The agent must add a Retry step to a workflow, connect it correctly, and move keyboard focus to that new step. The test checks the real browser, not just the agent's final answer. What happened We ran each version 20 times with gpt-5-mini using low reasoning effort. Tool result Verified… See the full description on the dataset page: https://huggingface.co/datasets/accesslint/2d-webmcp-browser-focus.tabularn<1K0 likes23 downloads1mo agoHugging Face16metehan777 /awesome-browser-use-prompts Awesome Browser-Use Prompts A curated collection of effective prompts for Browser-Use, the framework that enables AI agents to control web browsers. This repository aims to provide examples, templates, and best practices for crafting prompts that maximize the capabilities of Browser-Use agents. Introduction Browser-Use allows language models to interact with web interfaces through natural language instructions. Effective prompting is crucial to achieving successful… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/awesome-browser-use-prompts.textn<1K4 likes18 downloads2y agoHugging Face17Laramie2 /browseragent-datatext10K<n<100K0 likes14 downloads3mo agoHugging Face18Georgefifth /tiny-browser-planner-reason-dataset language: en task_categories: text-generation tags: reasoning planning browser-agent build-small size_categories: n<1K pretty_name: TinyBrowserPlanner Reason Dataset TinyBrowserPlanner-Reason Dataset Dataset used to train the Reason-First version of TinyBrowserPlanner. Key Result 4/12 → 10/12 planning accuracy Contains reasoning examples replanning scenarios wrong-page recovery paywall recovery refine_search examples… See the full description on the dataset page: https://huggingface.co/datasets/Georgefifth/tiny-browser-planner-reason-dataset.textn<1K0 likes13 downloads4mo agoHugging Face19Roy229 /probe-browser-evaltextn<1K0 likes11 downloads2mo agoHugging Face20Vendex /minicpm5-computer-browser-coding-v2textn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.