datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webshop-dataessential-web-v1.0
🌐 Essential-Web: Complete 24-Trillion Token Dataset
🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS
📋 Dataset Description
Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents.
Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.infinity-mm-stage1-webdataset
WebDataset Image-Text Dataset
This dataset contains image-text pairs in WebDataset format.
Dataset Structure
Each .tar.wds file contains entries with JSON data including image, text, and metadata.
open-web-math
Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba
GitHub | ArXiv
| PDF
OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models.
You can download the dataset using Hugging Face:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-web-math/open-web-math.weblinx-browsergym
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the video tag.
This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information.
[!NOTE]
The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.webtoepub-libraryCorpus-200B
WebOrganizer/Corpus-200B
[Paper] [Website] [GitHub]
This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with
(1) RefinedWeb filters and
(2) BFF deduplication.
We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores.
Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset().
The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.WebVidwaymo_webdatasetWebSight
Dataset Card for WebSight
Dataset Description
WebSight is a large synthetic dataset containing HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot.
This dataset serves as a valuable resource for tasks such as generating UI codes from a screenshot.
It comes in two versions:
v0.1: Websites are coded with HTML + CSS. They do not include real images.
v0.2: Websites are coded with HTML + Tailwind CSS. They do… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/WebSight.homelab-webdev
homelab-webdev
Private SFT mix for Gemma web-dev training. Training JSONL lives in dataset.merged.jsonl.
Public source datasets (catalog)
These Hugging Face datasets are the corpora we stream from (subsets only; full dumps are huge):
Key
Hub ID
Notes
websight
HuggingFaceM4/WebSight (v0.2)
Screenshot → HTML/Tailwind
webcode2m
xcodemind/webcode2m_purified
Real-world pages
design2code
SALT-NLP/Design2Code
~485 eval/VLM pages
stack_v2… See the full description on the dataset page: https://huggingface.co/datasets/swingdoor45/homelab-webdev.web_questions
Dataset Card for "web_questions"
Dataset Summary
This dataset consists of 6,642 question/answer pairs.
The questions are supposed to be answerable by Freebase, a large knowledge graph.
The questions are mostly centered around a single named entity.
The questions are popular ones asked on the web (at least in 2013).
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.swiss-caselaw-web-ui
Swiss Case Law Open Dataset
962,724 published decisions from Swiss federal, cantonal, and regulatory bodies.
Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh.
What this is
A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.webtoepub-library
LinkToEpub Community Library
EPUBs converted on https://linktoepub.com. Older books live in prasadonly/webtoepub-library.
webui-350kThis data accompanies the WebUI project (https://dl.acm.org/doi/abs/10.1145/3544548.3581158)
For more information, check out the project website: https://uimodeling.github.io/
To download this dataset, you need to install the huggingface-hub package
pip install huggingface-hub
Use snapshot_download
from huggingface_hub import snapshot_download
snapshot_download(repo_id="biglab/webui-350k", repo_type="dataset")
IMPORTANT
Before downloading and using, please review the copyright info here:… See the full description on the dataset page: https://huggingface.co/datasets/biglab/webui-350k.essential-web-1t-sample-fdc-partitioned
🌐 Essential-Web: FDC Level-2 Partitioned Dataset
📋 Dataset Description
This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering.
🔍 Free Decimal Correspondence (FDC)
The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.webapps
Neurodesk webapps assets
Models and validation data for https://github.com/neurodesk/webapps, which stores source only.
Layout and naming:
<app>/ — one directory per app, named as in apps/<app> or exes/<app>; shared assets live under the app that owns them.
<app>/models/ — weights. Upstream artifacts keep their upstream filename verbatim (synthsr_v20_230130.h5); derived artifacts are <app>-<variant>.onnx (synthsr-v2.onnx, synthstrip-browser.onnx). Precision or graph edits get a… See the full description on the dataset page: https://huggingface.co/datasets/neurodeskorg/webapps.webfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net
Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch.
Collected as part of my work on web simulators.
https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web.
https://x.com/leeleepenkman
Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.web-crawl-2026
Web Crawl 2026
A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project.
Dataset Description
This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped.
Data Format
Each record is a JSON line (gzipped) with fields:
text: extracted text content (200-200,000 chars)
url: source URL
domain: source domain
timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.conceptual-captions-12m-webdataset-bertswebvid-10Mwebagentscocis-web-info
COCIS WEB INFO
Dataset Summary
This dataset contains information about Makerere University College of Computing and Information Science that was scraped from its official website and corresponding websites.
The dataset consists of approximately 513 JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks.
By sharding the data into 513 files, this repository supports… See the full description on the dataset page: https://huggingface.co/datasets/jimjunior/cocis-web-info.pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2.
Dataloading code can be found here.
webis-touche2020-v3
Touche2020Retrieval.v3
An MTEB dataset
Massive Text Embedding Benchmark
Touché Task 1: Argument Retrieval for Controversial Questions
Task category
t2t
Domains
Academic
Reference
https://github.com/castorini/touche-error-analysis
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["Touche2020Retrieval.v3"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/webis-touche2020-v3.webui-testThis data accompanies the WebUI project (https://dl.acm.org/doi/abs/10.1145/3544548.3581158)
For more information, check out the project website: https://uimodeling.github.io/
To download this dataset, you need to install the huggingface-hub package
pip install huggingface-hub
Use snapshot_download
from huggingface_hub import snapshot_download
snapshot_download(repo_id="biglab/webui-test", repo_type="dataset")
IMPORTANT
Before downloading and using, please review the copyright info here:… See the full description on the dataset page: https://huggingface.co/datasets/biglab/webui-test.webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
Features:
image: the screenshot of the webpage.
bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information.
text: the webpage code text including HTML/CSS code.
scale: the scale of the screenshot, in the format [width, height].
lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.conceptual-captions-12m-webdataset-metadata
Conceptual Captions 12M — Webshart metadata indices
Per-shard webshart metadata indices for
laion/conceptual-captions-12m-webdataset:
1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout.
Each index records every tar member's byte offset and length (enabling ranged reads without
downloading whole shards), image geometry (width/height for aspect bucketing), and — as of
August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.
