Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TempoFunk /webvid-10Mtexttext-to-video10M<n<100M98 likes6.2k downloads3y agoHugging Face02APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes876 downloads16d agoHugging Face030xscope /web3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models. text1M<n<10M8 likes688 downloads2y agoHugging Face04BeIR /webis-touche2020-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.tabulartext-retrieval1K<n<10K0 likes569 downloads4y agoHugging Face05DeusHorizon /agent-web-index Agent Web Index — how much of the web can AI assistants actually read? 50,413 domains measured live. 26% of them cannot be read by at least one of ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/ Every row here is the result of real HTTP requests, not an estimate and not a re-publication of someone else's crawl: each domain's homepage is requested once as a browser and once as each of the published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.tabular10K<n<100K0 likes557 downloads21h agoHugging Face06bhadresh-savani /web_splitWork In progress! text1M<n<10M1 likes192 downloads5y agoHugging Face07snncn /shopify-websites Shopify Websites Dataset Dataset Description The Shopify Websites Dataset is a curated collection of 10,000 verified Shopify-powered e-commerce store URLs, providing researchers and analysts with a comprehensive resource for studying the Shopify e-commerce ecosystem. This dataset offers a diverse snapshot of real-world Shopify stores across various industries and geographies, making it valuable for market research, web scraping projects, and e-commerce platform analysis.… See the full description on the dataset page: https://huggingface.co/datasets/snncn/shopify-websites.text10K<n<100K1 likes189 downloads11mo agoHugging Face08Pijush22049 /casia-webfaceimage100K<n<1M0 likes175 downloads5mo agoHugging Face09river-martin /web-of-science-with-label-texts Dataset Description: The data is partitioned according to a 75/15/15 train/test/validate split. Each entry has an abstract (which is the input text for classification), a domain (a label from the list below), and an area (a subdomain of the paper, such as CS -> computer graphics, which takes on one of 134 possible values). All the attributes are strings. Domain labels: - Computer Science - Electrical Engineering - Psychology - Mechanical Engineering, - Civil Engineering - Medical… See the full description on the dataset page: https://huggingface.co/datasets/river-martin/web-of-science-with-label-texts.text10K<n<100K1 likes151 downloads2y agoHugging Face10ST-WebAgentBench /st-webagentbenchgated A Benchmark for Evaluating Safety & Trustworthiness in Web Agents Accepted at ICLR 2026 Overview ST-WebAgentBench is a policy-enriched evaluation suite for web agents, built on BrowserGym. It measures not only whether agents complete tasks, but whether they do so while respecting safety and trustworthiness (ST) policies — the constraints that govern real enterprise deployments. The… See the full description on the dataset page: https://huggingface.co/datasets/ST-WebAgentBench/st-webagentbench.textother1K<n<10K5 likes125 downloads7mo agoHugging Face11mindweave /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/web-server-logs.tabulartabular-classification1K<n<10K0 likes124 downloads6mo agoHugging Face12Halluminate /WebBench Web Bench: A real-world benchmark for Browser Agents WebBench is an open, task-oriented benchmark that measures how well browser agents handle realistic web workflows. It contains 2 ,454 tasks spread across 452 live websites selected from the global top-1000 by traffic. Last updated: May 28, 2025 Dataset Composition Category Description Example Count (% of dataset) READ Tasks that require searching and extracting information “Navigate to the news section and… See the full description on the dataset page: https://huggingface.co/datasets/Halluminate/WebBench.text1K<n<10K15 likes121 downloads1y agoHugging Face13lucas-ventura /WebVid-CoVRarxiv.org/abs/2308.14746 tabular1M<n<10M5 likes108 downloads2y agoHugging Face14UI-Simulator /UI-Simulator_web_datatext10K<n<100K1 likes106 downloads1y agoHugging Face15FirstBML1 /afrofinchain-multilingual-web3 AfroFinChain — Multilingual Web3 & Blockchain Dataset Multilingual Web3 & blockchain dataset in Yoruba, Hausa, Igbo, and Nigerian Pidgin with 1,451 terminology entries and 1,451 conversational Q&A pairs. Designed for LLM fine-tuning, financial literacy, and conversational AI in low-resource African languages. Uses culturally grounded analogies (e.g., ajo, adashi, isusu) to make DeFi concepts actually understandable. Built with Adaptive Data by Adaption as part of the Adaption… See the full description on the dataset page: https://huggingface.co/datasets/FirstBML1/afrofinchain-multilingual-web3.texttext-generation1K<n<10K0 likes87 downloads5mo agoHugging Face16shengqin /web-attacks-longtabular10K<n<100K4 likes76 downloads3y agoHugging Face174verburga /peru-web-tech-stacks Motivation I always though we were lacking common jargon datasets on tech detection. As an engineer I talk about fullstack frameworks or frontend frameworks, not about standalone technologies with low information about the architecture of the app. Next.js, Vue, React SPA, rather than React Router, core-js, Strapi CDN, and so forth. Content This dataset contains human-verified tech stack labels for Peruvian websites based on wappalyzer and manual inspection.… See the full description on the dataset page: https://huggingface.co/datasets/4verburga/peru-web-tech-stacks.texttabular-classificationn<1K0 likes75 downloads16d agoHugging Face18shengqin /web-attackstabular10K<n<100K12 likes74 downloads3y agoHugging Face19Plugiloinc /45_Million_Websitestabular1M<n<10M1 likes72 downloads1y agoHugging Face20convergence-ai /WebVoyager2025Valid WebVoyager 2025 Valid A modified subset of WebVoyager designed to be valid until 20th December 2025. This was used to benchmark proxy-lite. You can find the original WebVoyager tasks here. textn<1K7 likes63 downloads2y agoHugging Face21webxos /nexus-fpv NEXUS FPV Physics Dataset Sampler by webXOS Auto-generated small FPV drone flight-telemetry and reinforcement-learning for experience sample dataset captured directly in the browser by NEXUS FPV. Each row is one physics tick recorded while a drone flew through a waypoint course, either under manual control or the built-in PID auto-pilot. Play the game and make your own datasets: https://webxos.itch.io/nexus-fpv or download it from the /gym/ folder of this repo. Generator: NEXUS… See the full description on the dataset page: https://huggingface.co/datasets/webxos/nexus-fpv.tabularreinforcement-learningn<1K1 likes63 downloads3mo agoHugging Face22DEplain /DEplain-web-doc DEplain-web-doc: A corpus for German Document Simplification DEplain-web-doc is a subcorpus of DEplain Stodden et al., 2023 for document simplification. The corpus consists of 396 (199/50/147) parallel documents crawled from the web in standard German and plain German (or easy-to-read German). All documents are either published under an open license or the copyright holders gave us the permission to share the data. If you are interested in a larger corpus, please check our paper… See the full description on the dataset page: https://huggingface.co/datasets/DEplain/DEplain-web-doc.tabularn<1K1 likes59 downloads3y agoHugging Face23Webopen2026 /enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports Dataset Statistics Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB Languages: English French Spanish Main fields: email_id thread_id timestamp language bank department country risk_level Enterprise Financial Crime AI Dataset The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.tabulartext-classification100K<n<1M0 likes57 downloads7mo agoHugging Face24hug-the-trees /0skeng-agentic-web-benchmark 0skeng Index: synthetic UK collectibles prices The prices in this dataset are generated, not collected. A script produces every figure from a fixed seed. No marketplace was read, no listing was matched, and no real sale lies behind any number here. Do not use it to value, buy or sell anything, and do not repeat any figure in it as a market fact. What it is good for is testing: it has the shape, spread and internal consistency of a real weekly price index, in formats a program or… See the full description on the dataset page: https://huggingface.co/datasets/hug-the-trees/0skeng-agentic-web-benchmark.tabularn<1K1 likes53 downloads3d agoHugging Face25webimmunization /COVID-19-vaccine-attitude-tweets Dataset Card for COVID-19-vaccine-attitude-tweets Dataset Summary The dataset consists of 2564 manually annotated tweets related to COVID-19 vaccines. The dataset can be used to discover the attitude expressed in the tweet towards the subject of COVID-19 vaccines. Tweets are in English. The dataset was curated in such a way as to maximize the likelihood of tweets with a strong emotional tone. We have assumed the existence of three classes: PRO (label 0): positive, the… See the full description on the dataset page: https://huggingface.co/datasets/webimmunization/COVID-19-vaccine-attitude-tweets.tabulartext-classification1K<n<10K2 likes50 downloads4y agoHugging Face26mp-cmd /swiss-sme-websites-2026 Swiss SME Websites 2026: 10,000 sites audited, canton by canton Aggregated measurements of 10,000 websites of Swiss small and medium-sized businesses, drawn at random from the Swiss commercial register in September 2026. Proportions only: no company name, no website address, no figure for a group of fewer than 30 sites. Study page (FR): https://hallebardier.com/etude-sites-pme-suisses/ Study page (EN): https://hallebardier.com/en/swiss-sme-website-study/ Study page (DE):… See the full description on the dataset page: https://huggingface.co/datasets/mp-cmd/swiss-sme-websites-2026.tabularn<1K0 likes49 downloads4d agoHugging Face27linaaaaaaaaaaaaa /web-server-logs Web Server Access Logs (Synthetic) (Free Sample) This is a free sample with 5,003 rows. The full dataset has 50,048 rows across 2 tables. Realistic HTTP access logs from a simulated SaaS company running an e-commerce API and marketing website. 50,000 requests across 3 servers over 12 months. Includes realistic patterns: weekday/weekend traffic variation, peak hours, seasonal trends, bot traffic, and two injected anomalies (DDoS attempt and database outage) for anomaly… See the full description on the dataset page: https://huggingface.co/datasets/linaaaaaaaaaaaaa/web-server-logs.tabulartabular-classification1K<n<10K1 likes47 downloads4mo agoHugging Face28YangYang-Research /web-attack-detectionThe dataset contains 625,904 attack payload samples, with 294,771 labeled as 1 and 331,129 labeled as 0, including SQL injection, XSS, command injection, and other vulnerabilities. tabulartext-classification100K<n<1M1 likes46 downloads2y agoHugging Face29OmriShtayer /Website_Traffic_and_Engagementtabulartable-question-answeringn<1K0 likes41 downloads2y agoHugging Face30WebSEM-ai /agent-discoverability-ado-score-romania Agent Discoverability (ADO Score) — Romania, September 2026 130 Romanian domains probed for A2A Agent Cards, MCP discovery, llms.txt, schema.org and Wikidata. Zero Agent Cards; mean ADO Score 17/100. Raw data, scripts and scoring spec, CC BY 4.0. Canonical study (analysis, charts, interpretation): Romanian · English What this is On 8 September 2026 a standard-library Python probe (published) requested, for each of 130 domains, the homepage without JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/WebSEM-ai/agent-discoverability-ado-score-romania.tabularn<1K0 likes40 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.