Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01webagentlab /webchain WebChain v2 A large-scale, human-annotated dataset of real-world web interaction trajectories for training and evaluating web agents. [Paper] [Code] [Dataset] WebChain captures how people complete real tasks on live websites. It is designed for agents that must both identify the correct interface element and reason through a sequence of actions. Each trajectory aligns screenshots, web structure, grounded actions, and reasoning signals instead of treating web navigation as… See the full description on the dataset page: https://huggingface.co/datasets/webagentlab/webchain.tabular1K<n<10K2 likes3k downloads1mo agoHugging Face02OctoThinker /MegaMath-Web-Pro-Max OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year; Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data; Step 3: Training a fasttext carefully with proper preprocessing; Step 4: Filtering documents with a threshold (i.e., 0.4); Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.tabular10M<n<100M41 likes2.5k downloads1y agoHugging Face03NoeFlandre /osm-polygon-website-tag OSM Polygon Website Dataset OpenStreetMap polygons that carry a website or contact:website tag, with the full main-page text of each site. Every number below is recomputed from the published Parquet files. At a glance Polygons 1,726,474 With extracted text 1,192,980 Words of text 407,685,655 Languages 397 Regional sources 386 / 386 Duplicate objects removed 104,927 Candidates rejected 868,905,743 Status Done Website… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag.tabular1M<n<10M2 likes1.5k downloads11d agoHugging Face04AdaMLLab /WebTerminal Terminal/CLI Web Text A filtered extract of terminal and command-line content from two large web-text corpora, designed for upsampling agentic-adjacent data during pretraining. Subsets Subset Rows Tokens Size Quality clean (default) 2.33M 4.6B 11 GB ~98% terminal content unfiltered 61.3M 359B 962 GB ~15% terminal content from datasets import load_dataset # Load the clean subset (default) ds = load_dataset("AdaMLLab/WebTerminal") # Load the unfiltered… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/WebTerminal.tabulartext-generation10M<n<100M4 likes1.5k downloads8mo agoHugging Face05webninjasi /pk-map-statstabular1M<n<10M3 likes1.1k downloads1mo agoHugging Face06placeholderlabs /pretrain-web-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 73,646,210,143 (73.6B) Trainable tokens 73,646,210,143 (73.6B) Documents 61,059,647 Shards 590 UTF-8 bytes 341,537,872,441 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix.tabular100M<n<1B0 likes1k downloads29d agoHugging Face07APProjects /us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered Rebuilt 2026-09-25. Five state labor agencies — Connecticut, Michigan, New York, North Carolina and Pennsylvania — retired the web pages their older WARN Act layoff notices lived on. Their current pages start years later. This dataset is every notice in our file that came from one of those retired pages and is not on the agency's live page today: 6,799 notices, 6,799… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.tabulartabular-classification1K<n<10K0 likes876 downloads16d agoHugging Face08HAERAE-HUB /KOREAN-WEBTEXT KOREAN-WEBTEXT KOREAN-WEBTEXT is a high-quality Korean language corpus consisting of 2.2 billion tokens. The data has been collected from the following sources: cc100 oscar-corpus/OSCAR-2201 oscar-corpus/OSCAR-2109 oscar-corpus/OSCAR-2301 ontocord/CulturaY Additional credible internet sources collected by out team (We are working to add more sources) The dataset undergoes rigorous filtering at both the sentence and document levels to ensure quality of text data. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KOREAN-WEBTEXT.tabular1M<n<10M49 likes711 downloads2y agoHugging Face09MinjaeLee-FuriosaAI-Ext /ai-research-berkeley-webagentgated Berkeley WebAgent Experiment Artifacts Native GEPA, CLUE and ACE experiment logs and available actor trajectory evidence. Files require manual access approval. Request access with your Hugging Face account. Results and complete evidence snapshot — September 23, 2026 Combined experiment summary: GEPA, CLUE and ACE; method × site for WebArena. Non-WebArena repeat scores and variance, including completed ACE ≤50k-token evaluations. WebArena / GoBrowse progress and… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-webagent.tabularn<1K0 likes641 downloads14d agoHugging Face10openai /webgpt_comparisonsWebGPT Comparisons contains all of the comparisons marked as suitable for reward modelling from the WebGPT paper.tabular10K<n<100K242 likes601 downloads4y agoHugging Face11BeIR /webis-touche2020-qrels Dataset Card for BEIR Benchmark Dataset Summary BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks: Fact-checking: FEVER, Climate-FEVER, SciFact Question-Answering: NQ, HotpotQA, FiQA-2018 Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus News Retrieval: TREC-NEWS, Robust04 Argument Retrieval: Touche-2020, ArguAna Duplicate Question Retrieval: Quora, CqaDupstack Citation-Prediction: SCIDOCS Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/webis-touche2020-qrels.tabulartext-retrieval1K<n<10K0 likes569 downloads4y agoHugging Face12DeusHorizon /agent-web-index Agent Web Index — how much of the web can AI assistants actually read? 50,413 domains measured live. 26% of them cannot be read by at least one of ChatGPT, Claude, Perplexity or Gemini. Updated daily. Live index: https://shop.lumnika.com/ai-readiness/ Every row here is the result of real HTTP requests, not an estimate and not a re-publication of someone else's crawl: each domain's homepage is requested once as a browser and once as each of the published AI crawler user-agents… See the full description on the dataset page: https://huggingface.co/datasets/DeusHorizon/agent-web-index.tabular10K<n<100K0 likes557 downloads21h agoHugging Face13placeholderlabs /pretrain-web-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 8,689,580,607 (8.7B) Trainable tokens 8,689,580,607 (8.7B) Documents 281,846 Shards 89 UTF-8 bytes 37,540,769,483 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-web-mix-long-context.tabular100K<n<1M0 likes549 downloads29d agoHugging Face14Brainquiver /general-web-fr-202608 General · Web · French · 2026-08 French pretraining text, built from the French portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 31,999,309 documents and 118,346,333,763 characters of French prose. Contents Config Documents Characters Upstream fineweb2-hq-fra_Latn 31,999,309 118,346,333,763 epfml/FineWeb2-HQ, fra_Latn The character count is exact.… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-fr-202608.tabulartext-generation10M<n<100M0 likes498 downloads1mo agoHugging Face15Web3Survivor /Survivor 📚 FinePDFs-Edu 350B+ of highly educational tokens from PDFs 📄 What is it? 📚 FinePDFs-Edu dataset consists of 350B+ tokens of educational PDFs filtered from 📄 FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb-Edu, we developed an educational quality classifier using annotations generated by Qwen3-235B-A22B-Instruct-2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the… See the full description on the dataset page: https://huggingface.co/datasets/Web3Survivor/Survivor.tabulartext-generation10M<n<100M2 likes477 downloads10mo agoHugging Face16Brainquiver /general-web-it-202608 General · Web · Italian · 2026-08 Italian pretraining text, built from the Italian portion of EPFL's FineWeb2-HQ, which is the high quality slice of FineWeb-2. Every document passes one character-level cleaner and a repetition filter. 21,065,052 documents and 66,158,573,443 characters of Italian prose. Contents Config Documents Characters Upstream fineweb2-hq-ita_Latn 21,065,052 66,158,573,443 epfml/FineWeb2-HQ, ita_Latn The character count is… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-web-it-202608.tabulartext-generation10M<n<100M0 likes429 downloads1mo agoHugging Face17saksornr /fineweb-10bt-weborganizer-sigir-dota-rag FineWeb-10BT with WebOrganizer Labels for DoTA-RAG Dataset · DoTA-RAG paper · Project page This dataset is a labeled version of the FineWeb-10BT web corpus used as the document collection in DoTA-RAG: Dynamic of Thought Aggregation RAG. It contains 14,868,862 English-language documents in one train split. Each document retains its FineWeb text and provenance fields and adds a predicted topic and document format from WebOrganizer. The published Parquet files total 30.7 GB to… See the full description on the dataset page: https://huggingface.co/datasets/saksornr/fineweb-10bt-weborganizer-sigir-dota-rag.tabulartext-retrieval10M<n<100M1 likes354 downloads11d agoHugging Face18TurkuNLP /WebDocumentDescriptors Data release for the paper Task-Agnostic Web Document Annotation with LLM-Generated Descriptors (forthcoming). The descriptors are generated via a task-agnostic data annotation pipeline described in the paper (link coming soon). This Hugging Face dataset repository contains 5 distinct datasets: a descriptor-annotated version of a 10 billion token (~15 million document) sample of FineWeb. The 800k label descriptor schema The 500k document sample of FineWeb used to develop the schema along… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/WebDocumentDescriptors.tabulartext-classification10M<n<100M1 likes345 downloads25d agoHugging Face19AmineHA /WebArena-Verified WebArena-Verified Dataset description WebArena-Verified is a curated benchmark dataset of web tasks designed for reproducible evaluation of web agents across multiple realistic websites. Sources GitHub repository: webarena-verified Original WebArena benchmark: webarena.dev Splits full: 812 rows hard: 258 rows Tasks per site Counts below are task counts grouped by category. Tasks with more than one site are grouped under multi-category… See the full description on the dataset page: https://huggingface.co/datasets/AmineHA/WebArena-Verified.tabular1K<n<10K2 likes327 downloads8mo agoHugging Face20Lexmount /WebJev WebJev Training web agents to complete real tasks on the live web To complete a task on a real website, a web agent must get a long chain of decisions right, from the first page to the final answer: what to do next, and which element to act on. WebJev captures these decisions along complete task trajectories on the live web. It contains 64,122 decisions from 3,858 tasks on 1,443 real-world websites. They cover every stage of a task: searching, navigating, filtering, filling in… See the full description on the dataset page: https://huggingface.co/datasets/Lexmount/WebJev.tabularmultiple-choice100K<n<1M1 likes322 downloads11d agoHugging Face21Ba2han /mogan-turkish-web-long moganai/mogan-turkish-web long filtered Turkish texts Source: moganai/mogan-turkish-web (config: default, revision: d773a0efd1b7daf72c7909c83dbd385e7e3564d7). Rows contain 4,000–16,000 characters and passed the iteration-5 Turkish language, repetition, glue-word, punctuation, SEO, and soft information-density filters. Selected rows: 2,805,010. Generated by process_hf_dataset.py. See summary.json for counts and thresholds. tabulartext-generation1M<n<10M0 likes312 downloads21d agoHugging Face22abhijitramesh /webgpu-bench-leaderboardtabularn<1K1 likes292 downloads5mo agoHugging Face23lehduong /megamath-web-protabular10M<n<100M1 likes272 downloads1y agoHugging Face24pierretokns /seeclick-web-commercial-mlx SeeClick Web Commercial Dataset (MLX-VLM Format) Commercial-use friendly GUI grounding dataset from SeeClick Web data. Apache 2.0 licensed - safe for commercial applications. Dataset Description This dataset contains ~20k examples for training Vision-Language Models to predict click coordinates given a screenshot and instruction. Derived from SeeClick Web crawled data (Apache 2.0). Key Features License: Apache 2.0 (commercial use allowed) Format: MLX-VLM… See the full description on the dataset page: https://huggingface.co/datasets/pierretokns/seeclick-web-commercial-mlx.imageimage-to-text10K<n<100K0 likes261 downloads9mo agoHugging Face25axon-rl /webshop_instructionstabular1K<n<10K0 likes255 downloads1y agoHugging Face26AsciiMAster /pl-web-graph-2026-09-14 Polish Web Domain Observations 2026-09-14 A curated snapshot of a .pl-focused domain crawler: DNS observations, host availability metadata, discovered URL references, and the crawl frontier. No HTML, page text, or website classifications are included. Observations accumulated over months, so September 14 dates the export itself while each row carries its own observation time. Coverage is whatever one crawler reached, and liveness holds as of the recorded timestamp. Rendered… See the full description on the dataset page: https://huggingface.co/datasets/AsciiMAster/pl-web-graph-2026-09-14.tabular100M<n<1B0 likes250 downloads26d agoHugging Face27guanfengliu /so101_main_bin_2cameras_webThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 53, "total_frames": 9844, "total_tasks": 1, "total_videos": 106, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:53" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/guanfengliu/so101_main_bin_2cameras_web.tabularrobotics10K<n<100K0 likes247 downloads1y agoHugging Face28crawlora-net /dead-web-commoncrawl Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026) &nbsp;·&nbsp; Hugging Face &nbsp;·&nbsp; Kaggle &nbsp;·&nbsp; License: CC BY 4.0 An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.tabular100M<n<1B1 likes235 downloads3mo agoHugging Face29lightonai /webis-touche2020-decontaminated webis-touche2020 (Decontaminated) A decontaminated version of the webis-touche2020 dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed. Decontamination methodology Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files): Pass 1: Exact hash matching All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/webis-touche2020-decontaminated.tabulartext-retrieval100K<n<1M0 likes232 downloads7mo agoHugging Face30NoeFlandre /osm-polygon-website-tag-eunis Website This dataset preserves the source rows and adds four nullable EUNIS fields to each polygon-bearing Parquet shard. The map uses a representative point for each polygon and 2-degree geographic bins to keep the static artifact small. Colors identify EUNIS codes; the complete distribution is listed below. Unlabeled polygons are shown in gray. Label distribution EUNIS code Label Polygons Share Q11 Raised bog 89,423 5.18% Q41 Alkaline, calcareous… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-website-tag-eunis.tabular1M<n<10M0 likes211 downloads8d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.