datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
c4-en-html-with-training_metadata_allc4-en-html-with-metadata-ppl-cleanFile list:
"c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz",
"c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset
This dataset is currently under development
stackv2_html_filteredHtmlRAG-train
Dataset Information
We release the training data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The training data is sampled from 5 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, and MuSiQue. You can access the original dataset here.
We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-train.Natural_Questions_HTML_ToyNatural_Questions_HTML_reduced_allhtml-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.c4-en-html-with-metadatacc-html-tokens-qwen3
CommonCrawl HTML, tokenised
Packed token ids for language-model pretraining on raw web markup (not extracted text).
tokens
15.18B
shards
77
tokenizer
Qwen/Qwen3-0.6B
dtype
uint32, flat, no header
document separator
EOS token id
preprocessing
light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page
Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source
WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.frontend-html-tailwind-js
This dataset is a remastered version of Reubencf/frontend-coding prepared using Adaption's Adaptive Data platform.
frontend_html_tailwind_js
This dataset contains pairs of user prompts and generated frontend code solutions using HTML, Tailwind CSS, and JavaScript. The samples cover a wide range of web development tasks, including landing pages, portfolios, ecommerce sites, and interactive components with smooth scrolling and animations. Each entry demonstrates practical… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-html-tailwind-js.ui-distill-html-648
ui-distill-html-648
648 single-file HTML UI components, generated by
Ornith-1.0-35B on a single
RTX 3060 12GB, paired with the build request that produced each one.
Built to fine-tune a 3B model into writing UI
(DogukanUrker/ui-distill-3b), but
it stands on its own — distill your own student from it.
The interesting part
The instructions in this dataset are not the prompts that generated the HTML.
The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.MultiRoundConvos-Code-JS-HTML-CSS-Pythonhtml-to-markdown-10000-pairs
HTML to Markdown 10,000 Pairs
This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion.
Structure
pairs/train.jsonl: 9,000 examples
pairs/validation.jsonl: 500 examples
pairs/test.jsonl: 500 examples
manifest.json: generation manifest and split counts
Each JSONL row includes:
id: example identifier
html: source HTML string
markdown: target Markdown string
metadata: per-example metadata
Example
{"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.geocities-prompt-html
GeoCities prompt → HTML — fine-tune
Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page
from a plain-language description. This repo holds the dataset and the
training scripts so the whole thing runs from one place.
What's in here
dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line.
train_geocities.py — training entrypoint (loads this JSONL format).
train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.alpaca_webgen_htmlhtml_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
HTML_CSS_CodeDataSet_100kgemini-3-flash-preview-standalone-html-1kThis dataset was made by querying Gemini 3 Flash Preview to train models on web development for games, websites, and web apps in html + inline css and javascript.
Prompt breakdown:
First 800 targeted at standalone html websites, games, and apps.
The last 200 prompts are focused on science (chemistry, physics, biology), math, philosophy, multilingual tasks, and complex reasoning tasks.
All prompts were generated by gemini 3 pro preview
Stats:
Cost: $ 4.73 (USD)
Total tokens (input +… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview-standalone-html-1k.html
📚 HTML Tags Complete Reference Dataset
A comprehensive, structured dataset of HTML tags with detailed attributes, examples, and metadata. Perfect for learning HTML, building educational tools, or training ML models on web development concepts.
📊 Dataset Description
This dataset contains 85+ HTML tags with rich metadata including tag purpose, category, attributes, usage examples, and more. Each entry provides complete information about an HTML tag, making it an ideal… See the full description on the dataset page: https://huggingface.co/datasets/BarBar288/html.html_alpacachunked-enwiki-ns0-20250301-enterprise-html20250301.jsonl contains preprocessed text chunks from the wikipedia 20250301 html dump.
dead_letter_queue.jsonl contains articles that couldn't be processed due to formatitng or missing tags.
en_redirection_map.pkl can be deserialize to avoid building redirection_map everytime preprocessing_wikipedia_html_dump runs.
alpaca_data_with_html_outputhtml_datasetvov_thegioi_htmlsmall-fable-html-dataset
Dataset Card for 4cast/small-fable-html-dataset
This dataset provides a few high-quality examples of non-thinking outputs by Claude Fable 5. They have been taken from various sources such as ultimateplay.com.
Dataset Details
Dataset Sources
Repository: ultimateplay.com, claude.ai, arena.ai
vov_suckhoe_htmlimage-to-htmlhtmlembed
