Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YWZBrandon /webshop-data4 likes94k downloads1y agoHugging Face02EssentialAI /essential-web-v1.0 🌐 Essential-Web: Complete 24-Trillion Token Dataset 🏆 Website | 🖥️ Code | 📖 Paper | ☁️ AWS 📋 Dataset Description Essential-Web is a 24-trillion-token web dataset with document-level metadata designed for flexible dataset curation. The dataset provides metadata including subject matter classification, web page type, content complexity, and document quality scores for each of the 23.6 billion documents. Researchers can filter and curate specialized datasets… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/essential-web-v1.0.10B<n<100B247 likes81k downloads1y agoHugging Face03yangyang857658468 /infinity-mm-stage1-webdataset WebDataset Image-Text Dataset This dataset contains image-text pairs in WebDataset format. Dataset Structure Each .tar.wds file contains entries with JSON data including image, text, and metadata. 0 likes68k downloads2y agoHugging Face04open-web-math /open-web-math Keiran Paster*, Marco Dos Santos*, Zhangir Azerbayev, Jimmy Ba GitHub | ArXiv | PDF OpenWebMath is a dataset containing the majority of the high-quality, mathematical text from the internet. It is filtered and extracted from over 200B HTML files on Common Crawl down to a set of 6.3 million documents containing a total of 14.7B tokens. OpenWebMath is intended for use in pretraining and finetuning large language models. You can download the dataset using Hugging Face: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/open-web-math/open-web-math.text1M<n<10M363 likes37k downloads3y agoHugging Face05McGill-NLP /weblinx-browsergym WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the video tag. This dataset was specifically created to allow WebLINX to be used inside the BrowserGym and Agentlab ecosystem. Please see the browsergym repository for more information. [!NOTE] The version associated with this library is WebLINX… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/weblinx-browsergym.image-to-text4 likes30k downloads2y agoHugging Face06McGill-NLP /WebLINX-full WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead: WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WebLINX: Real-World Website Navigation with Multi-Turn Dialogue Xing Han Lù*, Zdeněk Kasner*, Siva Reddy 💾Code 📄Paper 🌐Website 📓Colab 🤖Models 💻Explorer 🐦Tweets 🏆Leaderboard Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.text10K<n<100K8 likes30k downloads1y agoHugging Face07prasadonly /webtoepub-library8 likes28k downloads14d agoHugging Face08WebOrganizer /Corpus-200B WebOrganizer/Corpus-200B [Paper] [Website] [GitHub] This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace's load_dataset(). The dataset has the following folder structure:… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/Corpus-200B.text100M<n<1B12 likes28k downloads3mo agoHugging Face09lucas-ventura /WebVidvideo1K<n<10K1 likes27k downloads12d agoHugging Face10Victer-XiaoyuYe /waymo_webdataset1 likes13k downloads1y agoHugging Face11HuggingFaceM4 /WebSight Dataset Card for WebSight Dataset Description WebSight is a large synthetic dataset containing HTML/CSS codes representing synthetically generated English websites, each accompanied by a corresponding screenshot. This dataset serves as a valuable resource for tasks such as generating UI codes from a screenshot. It comes in two versions: v0.1: Websites are coded with HTML + CSS. They do not include real images. v0.2: Websites are coded with HTML + Tailwind CSS. They do… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/WebSight.image1M<n<10M400 likes13k downloads3y agoHugging Face12swingdoor45 /homelab-webdev homelab-webdev Private SFT mix for Gemma web-dev training. Training JSONL lives in dataset.merged.jsonl. Public source datasets (catalog) These Hugging Face datasets are the corpora we stream from (subsets only; full dumps are huge): Key Hub ID Notes websight HuggingFaceM4/WebSight (v0.2) Screenshot → HTML/Tailwind webcode2m xcodemind/webcode2m_purified Real-world pages design2code SALT-NLP/Design2Code ~485 eval/VLM pages stack_v2… See the full description on the dataset page: https://huggingface.co/datasets/swingdoor45/homelab-webdev.text-generation0 likes12k downloads4d agoHugging Face13stanfordnlp /web_questions Dataset Card for "web_questions" Dataset Summary This dataset consists of 6,642 question/answer pairs. The questions are supposed to be answerable by Freebase, a large knowledge graph. The questions are mostly centered around a single named entity. The questions are popular ones asked on the web (at least in 2013). Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data… See the full description on the dataset page: https://huggingface.co/datasets/stanfordnlp/web_questions.textquestion-answering1K<n<10K43 likes11k downloads3y agoHugging Face14ArneH /swiss-caselaw-web-ui Swiss Case Law Open Dataset 962,724 published decisions from Swiss federal, cantonal, and regulatory bodies. Full text, structured metadata, and daily updates. The March 20, 2026 snapshot contains German, French, and Italian decisions; the export schema also reserves rm for Romansh. What this is A structured, searchable archive of Swiss court decisions — from the Federal Supreme Court (BGer) down to cantonal courts in all 26 cantons. Every decision includes the full… See the full description on the dataset page: https://huggingface.co/datasets/ArneH/swiss-caselaw-web-ui.0 likes11k downloads6mo agoHugging Face15prasaduser /webtoepub-library LinkToEpub Community Library EPUBs converted on https://linktoepub.com. Older books live in prasadonly/webtoepub-library. 0 likes10k downloads54m agoHugging Face16biglab /webui-350kThis data accompanies the WebUI project (https://dl.acm.org/doi/abs/10.1145/3544548.3581158) For more information, check out the project website: https://uimodeling.github.io/ To download this dataset, you need to install the huggingface-hub package pip install huggingface-hub Use snapshot_download from huggingface_hub import snapshot_download snapshot_download(repo_id="biglab/webui-350k", repo_type="dataset") IMPORTANT Before downloading and using, please review the copyright info here:… See the full description on the dataset page: https://huggingface.co/datasets/biglab/webui-350k.8 likes9.4k downloads3y agoHugging Face17Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8.7k downloads1y agoHugging Face18neurodeskorg /webapps Neurodesk webapps assets Models and validation data for https://github.com/neurodesk/webapps, which stores source only. Layout and naming: <app>/ — one directory per app, named as in apps/<app> or exes/<app>; shared assets live under the app that owns them. <app>/models/ — weights. Upstream artifacts keep their upstream filename verbatim (synthsr_v20_230130.h5); derived artifacts are <app>-<variant>.onnx (synthsr-v2.onnx, synthstrip-browser.onnx). Precision or graph edits get a… See the full description on the dataset page: https://huggingface.co/datasets/neurodeskorg/webapps.0 likes8.7k downloads20h agoHugging Face19lee101 /webfiddle-internet-raw-cache-datasetA dataset of different files that robots tried to crawl through webfiddle.net Mostly html files but other files too pdfs, images, binary- i have no idea what is in here at this stage - but gives an interesting idea of what crawlers like to visit and could be the basis of interesting SEO or coding LLM reasearch. Collected as part of my work on web simulators. https://webfiddle.net JS/CSS editor for the web, https://websim.netwrck.com Coding Editor for the web. https://x.com/leeleepenkman Its… See the full description on the dataset page: https://huggingface.co/datasets/lee101/webfiddle-internet-raw-cache-dataset.2 likes7.6k downloads47m agoHugging Face20OpenTransformer /web-crawl-2026 Web Crawl 2026 A large-scale web crawl dataset for language model pretraining, collected by the OpenTransformer project. Dataset Description This dataset contains text extracted from web pages crawled directly from the internet using custom high-throughput crawlers. All data is freshly scraped. Data Format Each record is a JSON line (gzipped) with fields: text: extracted text content (200-200,000 chars) url: source URL domain: source domain timestamp: crawl… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-2026.text-generation10B<n<100B1 likes7.6k downloads6mo agoHugging Face21AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes6.8k downloads2mo agoHugging Face22TempoFunk /webvid-10Mtexttext-to-video10M<n<100M98 likes6.2k downloads3y agoHugging Face23farama-minari /webagents0 likes5.7k downloads9mo agoHugging Face24jimjunior /cocis-web-info COCIS WEB INFO Dataset Summary This dataset contains information about Makerere University College of Computing and Information Science that was scraped from its official website and corresponding websites. The dataset consists of approximately 513 JSON chunks, designed for high-performance streaming and parallel processing. Each chunk represents a discrete unit of data structured for machine learning tasks. By sharding the data into 513 files, this repository supports… See the full description on the dataset page: https://huggingface.co/datasets/jimjunior/cocis-web-info.textquestion-answeringn<1K1 likes5.5k downloads6mo agoHugging Face25sayakpaul /pickapic_v2_webdatasetwebdataset archive of yuvalkirstain/pickapic_v2. Dataloading code can be found here. image1K<n<10K2 likes5.4k downloads3y agoHugging Face26mteb /webis-touche2020-v3 Touche2020Retrieval.v3 An MTEB dataset Massive Text Embedding Benchmark Touché Task 1: Argument Retrieval for Controversial Questions Task category t2t Domains Academic Reference https://github.com/castorini/touche-error-analysis How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["Touche2020Retrieval.v3"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/webis-touche2020-v3.texttext-retrieval100K<n<1M0 likes5.3k downloads1y agoHugging Face27biglab /webui-testThis data accompanies the WebUI project (https://dl.acm.org/doi/abs/10.1145/3544548.3581158) For more information, check out the project website: https://uimodeling.github.io/ To download this dataset, you need to install the huggingface-hub package pip install huggingface-hub Use snapshot_download from huggingface_hub import snapshot_download snapshot_download(repo_id="biglab/webui-test", repo_type="dataset") IMPORTANT Before downloading and using, please review the copyright info here:… See the full description on the dataset page: https://huggingface.co/datasets/biglab/webui-test.0 likes5.2k downloads3y agoHugging Face28xcodemind /webcode2m_purifiedWebCode2M: A Real-World Dataset for Code Generation from Webpage Designs Features: image: the screenshot of the webpage. bbox: the layout information, i.e., the bounding boxes (Bbox) of all the elements in the webpage, which contains the size, position, and hierarchy information. text: the webpage code text including HTML/CSS code. scale: the scale of the screenshot, in the format [width, height]. lang: the main language of the text content displayed on the rendered page (excluding HTML/CSS… See the full description on the dataset page: https://huggingface.co/datasets/xcodemind/webcode2m_purified.imageimage-to-text1M<n<10M6 likes5.1k downloads2y agoHugging Face29microsoft /webgym_tasks WebGym Tasks Dataset Dataset Description This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata. Dataset Summary Total Training Tasks: 292,092 Total Test Tasks: 1,167 Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.textreinforcement-learning100K<n<1M20 likes5k downloads8mo agoHugging Face30webshart /conceptual-captions-12m-webdataset-metadata Conceptual Captions 12M — Webshart metadata indices Per-shard webshart metadata indices for laion/conceptual-captions-12m-webdataset: 1,100 JSON files under data/, one per source tar shard, mirroring the source's shard layout. Each index records every tar member's byte offset and length (enabling ranged reads without downloading whole shards), image geometry (width/height for aspect bucketing), and — as of August 2026 — embedded captions for all 10,994,853 samples, coalesced… See the full description on the dataset page: https://huggingface.co/datasets/webshart/conceptual-captions-12m-webdataset-metadata.1 likes4.7k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.