Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /webgym_tasks WebGym Tasks Dataset Dataset Description This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata. Dataset Summary Total Training Tasks: 292,092 Total Test Tasks: 1,167 Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.textreinforcement-learning100K<n<1M20 likes4.9k downloads8mo agoHugging Face02wdndev /webnovel-chinese 简介 搜集网络上的网文小说,清洗,分割后,用于训练大语言模型,共计9000本左右,大约9B左右token。 使用 格式说明 采用jsonl格式存储,分为三个字段: title :小说名称 chapter:章节 text:正文内容 示例: {"title": "斗破苍穹", "chapter": " 第一章 陨落的天才", "text": "“斗之力,三段!”\n望着测验魔石碑上面闪亮得甚至有些刺眼的五个大字,少年面无表情,唇角有着一抹自嘲,紧握的手掌,因为大力,而导致略微尖锐的指甲深深的刺进了掌心之中,带来一阵阵钻心的疼痛……\n“萧炎,斗之力,三段!级别:低级!”测验魔石碑之旁,一位中年男子,看了一眼碑上所显示出来的信息,语气漠然的将之公布了出来……\n"} texttext-generation100K<n<1M48 likes3.3k downloads3y agoHugging Face03PaDaS-Lab /webfaq-retrievalWebFAQ Retrieval Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Retrieval Dataset is a carefully filtered and curated subset of the broader WebFAQ Q&A Dataset.It is purpose-built for Information Retrieval (IR) tasks, such as training and evaluating dense or sparse retrieval models in multiple languages. Each of the… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-retrieval.texttext-retrieval10M<n<100M10 likes3k downloads1y agoHugging Face04callanwu /WebWalkerQA📑 The paper of WebWalkerQA is available at arXiv. 📊 The dataset resource is a collection of 680 questions and answers from the WebWebWalker dataset. 🙋 The dataset is in the form of a JSON file. The keys in the JSON include: Question, Answer, Root_Url, and Info. The Info field contains more detailed information, including Hop, Domain, Language, Difficulty_Level, Source Website, and Golden_Path. { "Question": "When is the paper submission deadline for the ACL 2025 Industry Track, and what… See the full description on the dataset page: https://huggingface.co/datasets/callanwu/WebWalkerQA.textquestion-answering10K<n<100K53 likes2.7k downloads1y agoHugging Face05NJU-LINK /WebCompass WebCompass A unified multimodal benchmark for evaluating LLMs' ability to generate, edit, and repair functional web pages. WebCompass spans three input modalities — text design documents, reference screenshots, and video demonstrations — and three task families — generation, editing, and repair. GitHub: NJU-LINK/WebCompass Project Page: nju-link.github.io/WebCompass Quick Start from datasets import load_dataset # Generation tasks (existing) ds_text =… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/WebCompass.imagetext-generationn<1K6 likes2k downloads5mo agoHugging Face06facebook /recycling_the_web Dataset Card for Recycling-The-Web Synthetic Data We release 44.4B tokens of high-quality, model-filtered synthetic texts obtained via our REcycling the Web with guIded REwrite (REWIRE) approach. The generation process involves taking all documents that are of moderate quality (i.e., having passed some rule-based filters), using an LLM (Llama-3.3-70B-Instruct) to identify the purpose of the text content, and then asking the LLM to come up with an improved document conditioned on… See the full description on the dataset page: https://huggingface.co/datasets/facebook/recycling_the_web.text10M<n<100M68 likes1.4k downloads1y agoHugging Face07PaDaS-Lab /webfaq-bitextsWebFAQ Bilingual Datasets (Bitexts) Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Bilingual Datasets (a.k.a. Bitexts) are derived from the WebFAQ Q&A Dataset, but instead of monolingual question-answer (QA) pairs, each entry here contains aligned QA pairs in two different languages. These alignments are created via… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq-bitexts.texttext-retrieval1M<n<10M3 likes1.1k downloads2y agoHugging Face08PaDaS-Lab /webfaqWebFAQ Q&A Dataset Overview | Details | Structure | Examples | Considerations | License | Citation | Contact | Acknowledgement Overview The WebFAQ Q&A Dataset is a broad-coverage corpus of 96 million natural question-answer (QA) pairs in 75 languages, gathered from FAQ pages on the web. It leverages structured schema.org FAQPage annotations, making it a unique resource for large-scale Question Answering… See the full description on the dataset page: https://huggingface.co/datasets/PaDaS-Lab/webfaq.textquestion-answering10M<n<100M24 likes1.1k downloads1y agoHugging Face09mteb /cqadupstack-webmasters CQADupstackWebmastersRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackWebmastersRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-webmasters.texttext-retrieval10K<n<100K0 likes997 downloads1y agoHugging Face10Ancci /webqa_large7z x imgs.7z.001 text10K<n<100K0 likes991 downloads10mo agoHugging Face11MBZUAI /Web2Code Dataset Details Our Web2Code instruction tuning dataset construction and instruction generation process involves four key components: (1) Creation of new webpage image-code pair data: We generated high-quality HTML webpage-code pairs following the CodeAlpaca prompt using GPT-3.5 and convert them into instruction-following data. (2) Refinement of existing webpage code generation data: We transform existing datasets including into an instruction-following data format similar to LLaVA… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Web2Code.imagevisual-question-answeringn<1K13 likes938 downloads2y agoHugging Face12aisingapore /WangchanLION-Web Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.texttext-generation10M<n<100M3 likes900 downloads1y agoHugging Face13bytedance-research /Web-Bench Web-Bench English | 中文 README 📖 Overview Web-Bench is a benchmark designed to evaluate the performance of LLMs in actual Web development. Web-Bench contains 50 projects, each consisting of 20 tasks with sequential dependencies. The tasks implement project features in sequence, simulating real-world human development workflows. When designing Web-Bench, we aim to cover the foundational elements of Web development: Web Standards and Web Frameworks. Given the scale and… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/Web-Bench.text1K<n<10K11 likes704 downloads1y agoHugging Face14Qwen /WebWorldData WebWorldData 🌐 Overview WebWorldData is a large-scale dataset of 1.06M web interaction trajectories collected from the open web, designed for training browser world models. It is the training data behind the WebWorld model series. Each trajectory consists of sequences of (state, action, next_state) transitions, where states are represented as A11y Trees extracted from real websites using Playwright. Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/Qwen/WebWorldData.texttext-generation100K<n<1M81 likes697 downloads5mo agoHugging Face15MinjaeLee-FuriosaAI-Ext /ai-research-berkeley-webagentgated Berkeley WebAgent Experiment Artifacts Native GEPA, CLUE and ACE experiment logs and available actor trajectory evidence. Files require manual access approval. Request access with your Hugging Face account. Results and complete evidence snapshot — September 23, 2026 Combined experiment summary: GEPA, CLUE and ACE; method × site for WebArena. Non-WebArena repeat scores and variance, including completed ACE ≤50k-token evaluations. WebArena / GoBrowse progress and… See the full description on the dataset page: https://huggingface.co/datasets/MinjaeLee-FuriosaAI-Ext/ai-research-berkeley-webagent.tabularn<1K0 likes641 downloads14d agoHugging Face16Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes542 downloads21m agoHugging Face17WenyiWU0111 /webvoyager_evaluation_datatextn<1K0 likes536 downloads1y agoHugging Face18zxbsmk /webnovel_cn 内容 包含从12560本网文提取的约21.7M条可用于训练小说生成的中文指令数据(novel_json_tokens512.zip)。下载链接:https://pan.baidu.com/s/1TorBMbrqxrn6odRF0PJBVw 提取码:jlh3 以及从中提取出的包含50k条数据的子集(novel_cn_token512_50k.json)。其中输入和输出都不多于 512 tokens。 样例 在原有小说文本基础上,依据下列五种指令生成数据。 其中,文本由小说中随机抽取的连续句子组成。 给定标题,直接生成简介。 给定标题和简介,生成开头。 给定简介和一段文本,生成后续文本。 给定标题和一段文本,生成后续文本。 给定一段文本,生成后续文本。 { "instruction":… See the full description on the dataset page: https://huggingface.co/datasets/zxbsmk/webnovel_cn.text10K<n<100K130 likes458 downloads3y agoHugging Face19u8nXq3zW /ShowUI-web-8kimage1K<n<10K0 likes448 downloads1y agoHugging Face20Alibaba-NLP /WebShaper WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization Github: https://github.com/Alibaba-NLP/WebAgent Paper: https://arxiv.org/pdf/2507.15061 TLTR WebShaper is a synthesized training dataset for information-seeking (IS) task. It is based on our proposed task formalization of IS, and synthesized by our Expander Agent. WebShaper would cover a broader range of task forms, reasoning structure, and diversified knowledge. Description… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/WebShaper.textn<1K26 likes443 downloads1y agoHugging Face21Aunderline /webnlg_startext1K<n<10K0 likes415 downloads4y agoHugging Face22lmarena-ai /webdev-arena-preference-10k WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details in the blog post. Dataset License Agreement This Agreement contains the terms and conditions that govern your access and use of the WebDev Arena Dataset (Arena Dataset). You may not use the Arena Dataset if you do not accept this Agreement. By clicking to accept, accessing the Arena Dataset, or both, you hereby agree to the terms of the… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/webdev-arena-preference-10k.text10K<n<100K20 likes395 downloads2y agoHugging Face23BILGEM-AI /BILGE-Synthetic-Web BILGE-Synthetic-Web Dataset BILGE-Synthetic-Web was created following the methodology presented in the Cosmopedia blog/article. All content was generated using a 27B-parameter model. Further details on the methodology are available at: 🔗 https://huggingface.co/blog/cosmopedia text1M<n<10M9 likes381 downloads11mo agoHugging Face24zonghanHZH /ShowUI-web-8k ShowUI-web-8K This dataset is a curated 8K-sample subset from the original ShowUI-web dataset, as mentioned in our paper. It contributes to the training of GUI grounding models, with a focus on realistic web user interfaces collected from diverse websites. Dataset Details Source: Sampled from ShowUI-web Domain: Web GUI screenshots Diversity: Covers a wide variety of website layouts and components Use case: GUI grounding pretraining for web environments… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/ShowUI-web-8k.imageimage-text-to-text1K<n<10K0 likes354 downloads1y agoHugging Face25suolyer /webqatext10K<n<100K41 likes329 downloads4y agoHugging Face26AndyZijianZhang /webdsh-images webdsh-images Disk images for the emulated machines webdsh offers. Why this exists v86 can run about a hundred and twenty-five machines in a browser, and every one of them is the same emulator with a different disk. What copy.sh/v86 has that a fork does not is a CDN with the disks on it: its own host, i.copy.sh, refuses browser requests from anywhere else — deliberately, and it is their bandwidth to protect. So webdsh's catalog was complete and its machines were… See the full description on the dataset page: https://huggingface.co/datasets/AndyZijianZhang/webdsh-images.geospatialn<1K0 likes310 downloads2mo agoHugging Face27PersonalAILab /AFM-WebAgent-SFT-Dataset Data Introduction This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.text1K<n<10K10 likes292 downloads1y agoHugging Face28chungimungi /webkb-texttext1K<n<10K0 likes282 downloads1y agoHugging Face29Aunderline /webnlgtext1K<n<10K0 likes281 downloads4y agoHugging Face30chaannwooff /korean-web2 Keural-web (Naver Search) 한국어 웹 텍스트 코퍼스. 네이버 검색 API를 통해 수집한 한국어 문서 데이터셋입니다. 데이터 수집 방법 수집 도구: 네이버 검색 Open API (webkr, blog, news 엔드포인트) 검색 키워드: 경제·기술·사회·문화·과학·의학·법학·예술 등 27,192개 한국어 키워드 본문 추출: trafilatura 라이브러리로 HTML에서 본문 추출 수집 기간: 2025년 ~ 2026년 필터링 라인 단위 + 문서 단위 2단계 필터링이 적용된 상태입니다. 라인 단위 제거 항목: UI/네비게이션 잔재 (로그인, 더보기, 공유 버튼 등) 날짜·메타 태그, 기자 바이라인 광고·도박·성인 키워드 한국어 없는 줄 (영문, 중문, 일문 등) 해시태그, URL, 저작권 문구 등 문서 단위 필터링: 최소 텍스트 길이 미달 문서 제거 반복 문구 비율 초과 문서… See the full description on the dataset page: https://huggingface.co/datasets/chaannwooff/korean-web2.text10M<n<100M1 likes280 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.