datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
weblinx-browsergym
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the video tag.
This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information.
[!NOTE]
The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.BrowserAgent-Data
BrowserAgent ChatML Dataset (SFT/RFT)
This dataset contains ChatML-style multi-turn dialogues for a browser agent task. The data is prepared as JSON Lines so it can be previewed directly with the Hugging Face Hub Data Visualizer and loaded with the datasets library.
Links
Paper
Github
Files
sft.jsonl — SFT split (one JSON object per line)
rft.jsonl — RFT split (one JSON object per line)
Schema
Each record is a JSON object containing:
messages:… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/BrowserAgent-Data.BrowserART
Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents
Paper PDF
Homepage
Github
This project contains the behavior dataset in BrowserART, a red teaming test suit tailored particularly for browser agents.
Abstract
For safety reasons, large language models (LLMs) are trained to refuse harmful user instructions, such as assisting dangerous activities. We study an open question in this work: Can the desired safety refusal, typically enforced in chat… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/BrowserART.omnimcp_stealth_browser_scraping_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_stealth_browser_scraping_teaser.browser-agent-phase1-sft-action-only
Browser Agent Phase 1 SFT Action-Only
What this is
Action-only step-level chat SFT data for browser-agent training.
Each example teaches the model to predict the next BrowserGym action from:
the original generation-time system prompt used for data collection
task goal and URL
short recent history
current observation text and diagnostics
Assistant targets contain only the next action.
Why this format
This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.browser-tasks
Browser tasks
Live-browser tasks both pseudogenerated from large models and also taken from only from the starting_url field of
Halluminate/BrowserBench.
Each source page was revisited live at 800×600. Up to three tasks were generated
from the current page observation: information, navigation, and interaction.
Historical prompts, results, screenshots, and ground-truth URLs were not used.
The canonical dataset contains all 876 task slots from 292 source pages. A row is
accepted… See the full description on the dataset page: https://huggingface.co/datasets/merve/browser-tasks.hermes-flight-recorder-browser-tool-calling-trajectories
Hermes Flight Recorder Browser Tool-Calling Trajectories
This dataset repository publishes the exact public-synthetic artifacts used by
the Qwen3-4B browser LoRA case study.
data/browser/flightrecorder_action_sft.jsonl: governed browser train view.
data/development_action_sft.jsonl: frozen multi-scope development file; the
evaluator selects the browser task scope.
data/sealed_final_action_sft.jsonl: original frozen multi-scope final file;
the evaluator selects the browser task… See the full description on the dataset page: https://huggingface.co/datasets/zwright/hermes-flight-recorder-browser-tool-calling-trajectories.omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.omnimcp_browser_playwright_stealth_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_playwright_stealth_teaser.omnimcp_browser_session_pool_rotator_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_session_pool_rotator_teaser.omnimcp_browser_turnstile_solver_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_turnstile_solver_teaser.agentic_browser_sandbox_guard_teaser
🚀 AI Safety - Agentic Computer Use & Browser-Sandbox Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 AI Safety - Agentic Computer Use & Browser-Sandbox Guard on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100% AST-Valid Python)
PyArrow… See the full description on the dataset page: https://huggingface.co/datasets/emgena/agentic_browser_sandbox_guard_teaser.LexBench-Browser
LexBench-Browser
LexBench-Browser is a public browser-agent dataset for evaluating agents on real-web workflows.
The v1.0 snapshot contains 210 no-login tasks across 107 distinct websites, with Chinese and
English instructions, task-level reference steps, key points, common mistakes, scoring rubrics,
and robustness tags.
Repository: https://github.com/lexmount/browseruse-agent-bench
Dataset page: https://huggingface.co/datasets/Lexmount/LexBench-Browser
Docs:… See the full description on the dataset page: https://huggingface.co/datasets/Lexmount/LexBench-Browser.betterwright-agentic-browser-50k
BetterWright Agentic Browser — 6,093-row stopped checkpoint
This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings.
Contents
5,971 BetterWright demonstrations and 122 Playwright demonstrations.
32 task domains and 21 browser feature categories.
Harness-shaped conversations… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/betterwright-agentic-browser-50k.mcp-browser-automation-v2
MCP Browser Automation Training Dataset v2
Training data for teaching Vision-Language Models (VLMs) to use Claude Chrome MCP tools for browser automation.
Version History
Version
Date
Changes
v2
2026-01-28
Fixed tool names to use proper MCP namespaced format (mcp__claude-in-chrome__*)
v1
2026-01-27
Initial release with generic tool names
What's New in v2
Critical Fix: All tool names now use the correct MCP namespaced format that Claude Code… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/mcp-browser-automation-v2.tool-calling-browser-agent-tasks
Dataset Card
Created by: DataCreator AI
Overview
Tool Calling for Agentic Tasks with Multi-Step Workflows contains 1,062 synthetic multi-turn conversations between a user and an AI assistant. The examples primarily focus on practical agentic tasks such as train ticket booking, dynamic form filling, and payment processing. It provides diverse scenarios including successful execution, context retrieval, tool integration, and failure recovery.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/tool-calling-browser-agent-tasks.browser-agent-phase1-sft-reasoning-action
Browser Agent Phase 1 SFT Reasoning+Action
What this is
Reasoning-plus-action step-level chat SFT data for browser-agent training.
Each example uses the original generation-time system prompt, then appends a short instruction to reason first and output the final action.
Assistant targets contain:
one <think>...</think> block
then one BrowserGym action
Why this format
This is an experimental variant for comparing whether explicit reasoning supervision helps or… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-reasoning-action.sakthai-coder-browser
SakThai Coder Browser
Part of the SakThai Model Family.
Dataset Summary
SakThai Coder Browser is a synthetic instruction-tuning dataset for coding assistants and browser agents. It provides multi-turn agent-style conversations with structured tool definitions and expected assistant tool calls, built for training small language models on tool use and grounded code/browser workflows.
Owner: Nanthasit
Format: Parquet
Rows: 247
Columns: messages, tools
Task:… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-coder-browser.
