datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.custom-html-galleryc4-en-html-with-training_metadata_allaudio-htmlpubtabnet-htmlgithub-code-html-css-2webui-react-htmlcssjs-8740
WebUI React + HTML/CSS/JS 8,740
Curated export from ronantakizawa/webui containing every row where framework = react, plus 4,000 additional rows where framework = vanilla.
Screenshots: 8,740
React / vanilla HTML-CSS-JS rows: 4,740 / 4,000
Unique sample IDs: 2,914
Train / validation / test: 7,501 / 456 / 783
Viewports: 2,914 desktop / 2,913 mobile / 2,913 tablet
Images are stored as real image files and verified with Pillow.
viewer.html is a self-contained, paginated local… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/webui-react-htmlcssjs-8740.github-code-html-css-1c4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
htmlUN_Sitemap_Multilingual_HTML_Corpus
Dataset Card for "UN Sitemap Multilingual HTML Corpus"
Update Time +8:00 2023-3-23 17:25:20
Dataset Summary
此数据集是从联合国网站提供的的sitemap中爬取的,包含了各种语言的HTML文件,并按照语言进行分类。数据集包含了不同语言的文章、新闻等联合国文本。数据集旨在为研究人员、学者和语言技术开发人员提供一个多语言文本集,可用于各种自然语言处理任务和应用。
数据集包括以下语言:汉语(zh)、英语(en)、西班牙语(ar)、俄语(ru)、西班牙语(es)、法语(fr)。
Dataset Structure
Data Instances
数据集文件大小: 约 14 GB
一个 'zh' 的例子如下:
{
'uuid': 'a154688c-b385-4d2a-bec7-f239f1397d21',
'url':… See the full description on the dataset page: https://huggingface.co/datasets/ranWang/UN_Sitemap_Multilingual_HTML_Corpus.Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset
This dataset is currently under development
html-sampleHtmlRAG-train
Dataset Information
We release the training data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The training data is sampled from 5 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, and MuSiQue. You can access the original dataset here.
We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-train.stackv2_html_filteredusajobs_html_metadata
USAJOBS HTML Metadata Dataset (2017-01 to 2026-03)
Dataset Description
The USAJOBS HTML Metadata Dataset is a companion to usajobs_coded, which contains data extracted by the Job Ad Analysis Toolkit, and the raw USAJOBS HTML files at usajobs.
In this metadata version, we present the temporal anchors, canonical agency / department codes, and raw HTML banner audit-trail strings for each posting (keyed by usajobsControlNumber). Canonical codes use the OPM 2-character… See the full description on the dataset page: https://huggingface.co/datasets/loyoladatamining/usajobs_html_metadata.cleaned-htmlNatural_Questions_HTML_Toytable-image-html-pairsSCPWiki_HTML-ArchivesNatural_Questions_HTML_reduced_allgithub-code-html-cssfintabnet-htmlGutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.HtmlRAG-test
Dataset Information
We release the test data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The test data is sampled from 6 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, MuSiQue, and ELI5. You can access the original dataset here.
We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-test.html-eval
HTML Generation Evaluation
Overview
To evaluate our model's HTML generation performance, we have constructed a comprehensive evaluation dataset covering diverse web development scenarios. This dataset includes:
Web Development: Landing pages, e-commerce sites, responsive layouts, and financial dashboards
3D Scene Design: WebGL scenes, Three.js applications, and interactive 3D models
Game Design: Browser-based games, interactive gameplay, and physics-based games
Physics… See the full description on the dataset page: https://huggingface.co/datasets/nex-agi/html-eval.c4-en-html-with-metadataMind2Web-HTML-cleaned-lite-with-desc_w_tao
