datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MINT-1T-HTML
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.pubtabnet-htmlcc-html-tokens-qwen3
CommonCrawl HTML, tokenised
Packed token ids for language-model pretraining on raw web markup (not extracted text).
tokens
15.18B
shards
77
tokenizer
Qwen/Qwen3-0.6B
dtype
uint32, flat, no header
document separator
EOS token id
preprocessing
light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page
Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source
WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.ui-distill-html-648
ui-distill-html-648
648 single-file HTML UI components, generated by
Ornith-1.0-35B on a single
RTX 3060 12GB, paired with the build request that produced each one.
Built to fine-tune a 3B model into writing UI
(DogukanUrker/ui-distill-3b), but
it stands on its own — distill your own student from it.
The interesting part
The instructions in this dataset are not the prompts that generated the HTML.
The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.geocities-prompt-html
GeoCities prompt → HTML — fine-tune
Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page
from a plain-language description. This repo holds the dataset and the
training scripts so the whole thing runs from one place.
What's in here
dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line.
train_geocities.py — training entrypoint (loads this JSONL format).
train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.dataviz-html-dataset
DataViz HTML Dashboard Dataset
100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
data/*.csv — CSV data files used by 15 dashboards
README.md — Dataset card
Columns
Column
Type
Description
filename
string
File name of the dashboard
title
string
Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
crawlerlm-html-to-json
CrawlerLM: HTML to JSON Extraction
A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML.
Dataset Description
This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains.
Key Features
447 examples in instruction-tuning chat format
Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.50k-HTML-PRETRAIN
50k-HTML-PRETRAIN
Pretrain-style pairs: an English site assignment and a complete HTML document that implements it.
Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty.
Split
split
n
train
57,617
10 rows have a PNG in screenshot / images/. The other rows have a null screenshot.
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.philosophy-culture-translations-html-csv
AI-Culture Philosophy and Culture Translations CSV + HTML Corpus
The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind.
This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.dataviz-charts-html-dataset
DataViz Individual Charts HTML Dataset
120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations.
Structure
data/train-00000-of-00001.parquet — Main dataset in Parquet format
README.md — Dataset card
Columns
Column
Type
Description
filename
string
treemap-tech-innovation.html
title
string
e.g. "Bar Chart — Midnight Galaxy"
chart_type
string
bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.html-compression
HTML Compression Dataset
This dataset is designed for fine-tuning language models to understand and generate compressed HTML notation.
Dataset Description
The dataset contains pairs of standard HTML and compressed HTML using a custom compression algorithm that:
Abbreviates common HTML tags (div → d, span → s, etc.)
Shortens attribute names (class → c, style → st, etc.)
Compresses CSS patterns (display: flex → d:f, etc.)
Removes unnecessary whitespace and comments… See the full description on the dataset page: https://huggingface.co/datasets/tkattkat101/html-compression.geo_html_200
GEO HTML 200 Dataset
A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research.
Features
Column
Description
doc_id
Unique document identifier
url
Source URL
cleaned_text
Parsed plain text content
cleaned_text_length
Character count
query
Associated search query
title
Document title
topic_tags
Topic classification
Usage
from datasets import load_dataset
ds = load_dataset("erv1n/geo_html_200")
