Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B123 likes637k downloads2y agoHugging Face02bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes2.2k downloads4y agoHugging Face03apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M24 likes1.3k downloads2y agoHugging Face04hardikg2907 /github-code-html-css-2text1M<n<10M0 likes1k downloads2y agoHugging Face05Reubencf /webui-react-htmlcssjs-8740 WebUI React + HTML/CSS/JS 8,740 Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla. Screenshots: 8,740 React / vanilla HTML-CSS-JS rows: 4,740 / 4,000 Unique sample IDs: 2,914 Train / validation / test: 7,501 / 456 / 783 Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet Images are stored as real image files and verified with Pillow. viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.imageimage-to-text1K<n<10K0 likes975 downloads2mo agoHugging Face06hardikg2907 /github-code-html-css-1text1M<n<10M2 likes852 downloads2y agoHugging Face07akazemian /audio-htmltabular10K<n<100K0 likes745 downloads1y agoHugging Face08paralym /mint-1t-html-images-gte6-sample Size: 6769158 images sampled from Mint-1t-html Criteria: Data entries with greater than or equal to 6 images (gte6) image1M<n<10M0 likes643 downloads2y agoHugging Face09masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes560 downloads4y agoHugging Face10ranWang /UN_Sitemap_Multilingual_HTML_Corpus Dataset Card for "UN Sitemap Multilingual HTML Corpus" Update Time +8:00 2023-3-23 17:25:20 Dataset Summary 此数据集是从联合国网站提供的的sitemap中爬取的,包含了各种语言的HTML文件,并按照语言进行分类。数据集包含了不同语言的文章、新闻等联合国文本。数据集旨在为研究人员、学者和语言技术开发人员提供一个多语言文本集,可用于各种自然语言处理任务和应用。 数据集包括以下语言:汉语(zh)、英语(en)、西班牙语(ar)、俄语(ru)、西班牙语(es)、法语(fr)。 Dataset Structure Data Instances 数据集文件大小: 约 14 GB 一个 'zh' 的例子如下: { 'uuid': 'a154688c-b385-4d2a-bec7-f239f1397d21', 'url':… See the full description on the dataset page: https://huggingface.co/datasets/ranWang/UN_Sitemap_Multilingual_HTML_Corpus.text100K<n<1M3 likes479 downloads3y agoHugging Face11SaulLu /Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset This dataset is currently under development text10K<n<100K0 likes415 downloads5y agoHugging Face12big-computer /html-sampleimagen<1K0 likes397 downloads2y agoHugging Face13common-pile /stackv2_html_filteredtext1M<n<10M3 likes316 downloads1y agoHugging Face14zstanjj /HtmlRAG-train Dataset Information We release the training data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The training data is sampled from 5 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, and MuSiQue. You can access the original dataset here. We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-train.textquestion-answering10K<n<100K1 likes294 downloads1y agoHugging Face15SaulLu /Natural_Questions_HTML_Toytextn<1K0 likes272 downloads5y agoHugging Face16krytonguard /wiki_html_setstext10M<n<100M0 likes265 downloads11mo agoHugging Face17SaulLu /Natural_Questions_HTML_reduced_alltext10K<n<100K4 likes240 downloads5y agoHugging Face18hardikg2907 /github-code-html-csstext100K<n<1M4 likes227 downloads2y agoHugging Face19cognaize /table-image-html-pairsimage10K<n<100K0 likes215 downloads7mo agoHugging Face20FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes187 downloads1y agoHugging Face21apoidea /fintabnet-htmlimage10K<n<100K9 likes177 downloads2y agoHugging Face22bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes153 downloads4y agoHugging Face23PatrickHaller /cc-html-tokens-qwen3 CommonCrawl HTML, tokenised Packed token ids for language-model pretraining on raw web markup (not extracted text). tokens 15.18B shards 77 tokenizer Qwen/Qwen3-0.6B dtype uint32, flat, no header document separator EOS token id preprocessing light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.texttext-generationn<1K0 likes138 downloads11d agoHugging Face24LangAGI-Lab /Mind2Web-HTML-cleaned-lite-with-desc_w_taotabular1K<n<10K1 likes117 downloads2y agoHugging Face25hardikg2907 /github-code-html-css-split-3text1M<n<10M1 likes114 downloads2y agoHugging Face26Reubencf /frontend-html-tailwind-js This dataset is a remastered version of Reubencf/frontend-coding prepared using Adaption's Adaptive Data platform. frontend_html_tailwind_js This dataset contains pairs of user prompts and generated frontend code solutions using HTML, Tailwind CSS, and JavaScript. The samples cover a wide range of web development tasks, including landing pages, portfolios, ecommerce sites, and interactive components with smooth scrolling and animations. Each entry demonstrates practical… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-html-tailwind-js.textn<1K1 likes105 downloads6mo agoHugging Face272packer /html_game_gentextn<1K2 likes102 downloads1y agoHugging Face28williambrach /html-query-text-HtmlRAG html-query-text-HtmlRAG Warning: This dataset is under development and its content is subject to change! This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning. 🚀 Supported Tasks This dataset is primarily designed for: HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output. Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.textfeature-extraction10K<n<100K0 likes89 downloads11mo agoHugging Face29DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K2 likes88 downloads3mo agoHugging Face30mdhasnainali /job-html-to-jsontext10K<n<100K2 likes78 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.