datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
html-table-reconstruction-benchmark
HTML Table Reconstruction Benchmark
This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation.
The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.cc-html-tokens-qwen3
CommonCrawl HTML, tokenised
Packed token ids for language-model pretraining on raw web markup (not extracted text).
tokens
15.18B
shards
77
tokenizer
Qwen/Qwen3-0.6B
dtype
uint32, flat, no header
document separator
EOS token id
preprocessing
light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page
Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source
WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.ui-distill-html-648
ui-distill-html-648
648 single-file HTML UI components, generated by
Ornith-1.0-35B on a single
RTX 3060 12GB, paired with the build request that produced each one.
Built to fine-tune a 3B model into writing UI
(DogukanUrker/ui-distill-3b), but
it stands on its own — distill your own student from it.
The interesting part
The instructions in this dataset are not the prompts that generated the HTML.
The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.geocities-prompt-html
GeoCities prompt → HTML — fine-tune
Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page
from a plain-language description. This repo holds the dataset and the
training scripts so the whole thing runs from one place.
What's in here
dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line.
train_geocities.py — training entrypoint (loads this JSONL format).
train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
