Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B115 likes811k downloads2y agoHugging Face02gradio /custom-html-gallery0 likes4.5k downloads3mo agoHugging Face03bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes2.4k downloads4y agoHugging Face04akazemian /audio-htmltabular10K<n<100K0 likes1.5k downloads1y agoHugging Face05apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M24 likes1.2k downloads2y agoHugging Face06hardikg2907 /github-code-html-css-2text1M<n<10M0 likes900 downloads2y agoHugging Face07Reubencf /webui-react-htmlcssjs-8740 WebUI React + HTML/CSS/JS 8,740 Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla. Screenshots: 8,740 React / vanilla HTML-CSS-JS rows: 4,740 / 4,000 Unique sample IDs: 2,914 Train / validation / test: 7,501 / 456 / 783 Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet Images are stored as real image files and verified with Pillow. viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.imageimage-to-text1K<n<10K0 likes872 downloads2mo agoHugging Face08hardikg2907 /github-code-html-css-1text1M<n<10M2 likes849 downloads2y agoHugging Face09masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes751 downloads4y agoHugging Face10paralym /mint-1t-html-images-gte6-sample Size: 6769158 images sampled from Mint-1t-html Criteria: Data entries with greater than or equal to 6 images (gte6) image1M<n<10M0 likes748 downloads2y agoHugging Face11Moyuu /html0 likes681 downloads2y agoHugging Face12ranWang /UN_Sitemap_Multilingual_HTML_Corpus Dataset Card for "UN Sitemap Multilingual HTML Corpus" Update Time +8:00 2023-3-23 17:25:20 Dataset Summary 此数据集是从联合国网站提供的的sitemap中爬取的,包含了各种语言的HTML文件,并按照语言进行分类。数据集包含了不同语言的文章、新闻等联合国文本。数据集旨在为研究人员、学者和语言技术开发人员提供一个多语言文本集,可用于各种自然语言处理任务和应用。 数据集包括以下语言:汉语(zh)、英语(en)、西班牙语(ar)、俄语(ru)、西班牙语(es)、法语(fr)。 Dataset Structure Data Instances 数据集文件大小: 约 14 GB 一个 'zh' 的例子如下: { 'uuid': 'a154688c-b385-4d2a-bec7-f239f1397d21', 'url':… See the full description on the dataset page: https://huggingface.co/datasets/ranWang/UN_Sitemap_Multilingual_HTML_Corpus.text100K<n<1M3 likes510 downloads3y agoHugging Face13SaulLu /Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset This dataset is currently under development text10K<n<100K0 likes425 downloads5y agoHugging Face14big-computer /html-sampleimagen<1K0 likes357 downloads2y agoHugging Face15zstanjj /HtmlRAG-train Dataset Information We release the training data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The training data is sampled from 5 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, and MuSiQue. You can access the original dataset here. We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-train.textquestion-answering10K<n<100K1 likes339 downloads1y agoHugging Face16common-pile /stackv2_html_filteredtext1M<n<10M3 likes317 downloads1y agoHugging Face17loyoladatamining /usajobs_html_metadata USAJOBS HTML Metadata Dataset (2017-01 to 2026-03) Dataset Description The USAJOBS HTML Metadata Dataset is a companion to usajobs_coded, which contains data extracted by the Job Ad Analysis Toolkit, and the raw USAJOBS HTML files at usajobs. In this metadata version, we present the temporal anchors, canonical agency / department codes, and raw HTML banner audit-trail strings for each posting (keyed by usajobsControlNumber). Canonical codes use the OPM 2-character… See the full description on the dataset page: https://huggingface.co/datasets/loyoladatamining/usajobs_html_metadata.0 likes311 downloads5mo agoHugging Face18placingholocaust /cleaned-html1 likes281 downloads2y agoHugging Face19SaulLu /Natural_Questions_HTML_Toytextn<1K0 likes270 downloads5y agoHugging Face20cognaize /table-image-html-pairsimage10K<n<100K0 likes263 downloads7mo agoHugging Face21AiAF /SCPWiki_HTML-Archives0 likes247 downloads2y agoHugging Face22SaulLu /Natural_Questions_HTML_reduced_alltext10K<n<100K4 likes246 downloads5y agoHugging Face23hardikg2907 /github-code-html-csstext100K<n<1M4 likes226 downloads2y agoHugging Face24apoidea /fintabnet-htmlimage10K<n<100K9 likes220 downloads2y agoHugging Face25FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes205 downloads1y agoHugging Face26sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes202 downloads5mo agoHugging Face27zstanjj /HtmlRAG-test Dataset Information We release the test data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The test data is sampled from 6 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, MuSiQue, and ELI5. You can access the original dataset here. We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-test.question-answering1K<n<10K0 likes199 downloads1y agoHugging Face28nex-agi /html-eval HTML Generation Evaluation Overview To evaluate our model's HTML generation performance, we have constructed a comprehensive evaluation dataset covering diverse web development scenarios. This dataset includes: Web Development: Landing pages, e-commerce sites, responsive layouts, and financial dashboards 3D Scene Design: WebGL scenes, Three.js applications, and interactive 3D models Game Design: Browser-based games, interactive gameplay, and physics-based games Physics… See the full description on the dataset page: https://huggingface.co/datasets/nex-agi/html-eval.8 likes193 downloads11mo agoHugging Face29bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes177 downloads4y agoHugging Face30LangAGI-Lab /Mind2Web-HTML-cleaned-lite-with-desc_w_taotabular1K<n<10K1 likes176 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.