datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webvid-10Mus-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.web3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models.
webis-touche2020-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.agent-web-index
Agent Web Index — how much of the web can AI assistants actually read?
50,413 domains measured live. 26% of them cannot be read by at least one of
ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/
Every row here is the result of real HTTP requests, not an estimate and not a re-publication of
someone else's crawl: each domain's homepage is requested once as a browser and once as each of the
published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.web_splitWork In progress!
shopify-websites
Shopify Websites Dataset
Dataset Description
The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem.
This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.casia-webfaceweb-of-science-with-label-texts
Dataset Description:
The data is partitioned according to a 75/15/15 train/test/validate split.
Each entry has an abstract (which is the input text for classification), a domain (a label from the list below), and an area (a subdomain of the paper, such as CS -> computer graphics, which takes on one of 134 possible values).
All the attributes are strings.
Domain labels:
- Computer Science
- Electrical Engineering
- Psychology
- Mechanical Engineering,
- Civil Engineering
- Medical… See the full description on the dataset page: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts.st-webagentbench
A Benchmark for Evaluating Safety & Trustworthiness in Web Agents
Accepted at ICLR 2026
Overview
ST-WebAgentBench is a policy-enriched evaluation suite for web agents, built on BrowserGym. It measures not only whether agents complete tasks, but whether they do so while respecting safety and trustworthiness (ST) policies — the constraints that govern real enterprise deployments.
The… See the full description on the dataset page: https://huggingface.co/datasets/ST-WebAgentBench/st-webagentbench.web-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.WebBench
Web Bench: A real-world benchmark for Browser Agents
WebBench is an open, task-oriented benchmark that measures how well browser agents handle realistic web workflows.
It contains 2 ,454 tasks spread across 452 live websites selected from the global top-1000 by traffic.
Last updated: May 28, 2025
Dataset Composition
Category
Description
Example
Count (% of dataset)
READ
Tasks that require searching and extracting information
“Navigate to the news section and… See the full description on the dataset page: https://huggingface.co/datasets/Halluminate/WebBench.WebVid-CoVRarxiv.org/abs/2308.14746
UI-Simulator_web_dataafrofinchain-multilingual-web3
AfroFinChain — Multilingual Web3 & Blockchain Dataset
Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable.
Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.web-attacks-longperu-web-tech-stacks
Motivation
I always though we were lacking common jargon datasets on tech detection.
As an engineer I talk about fullstack frameworks or frontend frameworks, not about standalone technologies with low information about the architecture of the app.
Next.js, Vue, React SPA, rather than React Router, core-js, Strapi CDN, and so forth.
Content
This dataset contains human-verified tech stack labels for Peruvian websites based on wappalyzer and manual inspection.… See the full description on the dataset page: https://huggingface.co/datasets/4verburga/peru-web-tech-stacks.web-attacks45_Million_WebsitesWebVoyager2025Valid
WebVoyager 2025 Valid
A modified subset of WebVoyager designed to be valid until 20th December 2025.
This was used to benchmark proxy-lite.
You can find the original WebVoyager tasks here.
nexus-fpv
NEXUS FPV Physics Dataset Sampler by webXOS
Auto-generated small FPV drone flight-telemetry and reinforcement-learning for experience sample dataset captured directly in the browser
by NEXUS FPV. Each row is one physics tick recorded while a drone flew through a waypoint course, either under manual control or the built-in PID auto-pilot.
Play the game and make your own datasets: https://webxos.itch.io/nexus-fpv or download it from the /gym/ folder of this repo.
Generator: NEXUS… See the full description on the dataset page: https://huggingface.co/datasets/webxos/nexus-fpv.DEplain-web-doc
DEplain-web-doc: A corpus for German Document Simplification
DEplain-web-doc is a subcorpus of DEplain Stodden et al., 2023 for document simplification.
The corpus consists of 396 (199/50/147) parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license or the copyright holders gave us the permission to share the data.
If you are interested in a larger corpus, please check our paper… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-doc.enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports
Dataset Statistics
Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB
Languages:
English
French
Spanish
Main fields:
email_id
thread_id
timestamp
language
bank
department
country
risk_level
Enterprise Financial Crime AI Dataset
The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.0skeng-agentic-web-benchmark
0skeng Index: synthetic UK collectibles prices
The prices in this dataset are generated, not collected. A script produces
every figure from a fixed seed. No marketplace was read, no listing was
matched, and no real sale lies behind any number here. Do not use it to value,
buy or sell anything, and do not repeat any figure in it as a market fact.
What it is good for is testing: it has the shape, spread and internal
consistency of a real weekly price index, in formats a program or… See the full description on the dataset page: https://huggingface.co/datasets/hug-the-trees/0skeng-agentic-web-benchmark.COVID-19-vaccine-attitude-tweets
Dataset Card for COVID-19-vaccine-attitude-tweets
Dataset Summary
The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes:
PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.swiss-sme-websites-2026
Swiss SME Websites 2026: 10,000 sites audited, canton by canton
Aggregated measurements of 10,000 websites of Swiss small and medium-sized
businesses, drawn at random from the Swiss commercial register in September
2026. Proportions only: no company name, no website address, no figure for a
group of fewer than 30 sites.
Study page (FR): https://hallebardier.com/etude-sites-pme-suisses/
Study page (EN): https://hallebardier.com/en/swiss-sme-website-study/
Study page (DE):… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/swiss-sme-websites-2026.web-server-logs
Web Server Access Logs (Synthetic) (Free Sample)
This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables.
Realistic HTTP access logs from a simulated SaaS company running an
e-commerce API and marketing website. 50,000 requests across 3 servers
over 12 months.
Includes realistic patterns: weekday/weekend traffic variation, peak hours,
seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and
database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.web-attack-detectionThe dataset contains 625,904 attack payload samples, with 294,771 labeled as 1 and 331,129 labeled as 0, including SQL injection, XSS, command injection, and other vulnerabilities.
Website_Traffic_and_Engagementagent-discoverability-ado-score-romania
Agent Discoverability (ADO Score) — Romania, September 2026
130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0.
Canonical study (analysis, charts, interpretation):
Romanian ·
English
What this is
On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.
