Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-HTML 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-HTML.textimage-to-text100M<n<1B123 likes637k downloads2y agoHugging Face02apoidea /pubtabnet-htmlimagevisual-question-answering100K<n<1M24 likes1.3k downloads2y agoHugging Face03PatrickHaller /cc-html-tokens-qwen3 CommonCrawl HTML, tokenised Packed token ids for language-model pretraining on raw web markup (not extracted text). tokens 15.18B shards 77 tokenizer Qwen/Qwen3-0.6B dtype uint32, flat, no header document separator EOS token id preprocessing light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.texttext-generationn<1K0 likes138 downloads11d agoHugging Face04DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K2 likes88 downloads3mo agoHugging Face05thekevinscott /geocities-prompt-html GeoCities prompt → HTML — fine-tune Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page from a plain-language description. This repo holds the dataset and the training scripts so the whole thing runs from one place. What's in here dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line. train_geocities.py — training entrypoint (loads this JSONL format). train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.texttext-generation1K<n<10K2 likes50 downloads4mo agoHugging Face06sanjaymalladi /dataviz-html-dataset DataViz HTML Dashboard Dataset 100 HTML dashboard files built with the dash-viz-kit library — 10 themes, 12 chart types (ApexCharts + ECharts), zero config, declarative HTML. Structure data/train-00000-of-00001.parquet — Main dataset in Parquet format data/*.csv — CSV data files used by 15 dashboards README.md — Dataset card Columns Column Type Description filename string File name of the dashboard title string Human-readable title… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-html-dataset.tabulartext-generationn<1K0 likes33 downloads3mo agoHugging Face07Boakpe /deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript. texttext-generationn<1K1 likes32 downloads9mo agoHugging Face08espsluar /crawlerlm-html-to-json CrawlerLM: HTML to JSON Extraction A synthetic instruction-tuning dataset for training language models to extract structured JSON from HTML. Dataset Description This dataset contains HTML paired with structured JSON extraction tasks in chat format. It's designed for fine-tuning small language models to perform structured data extraction from messy, real-world HTML across multiple domains. Key Features 447 examples in instruction-tuning chat format Real HTML… See the full description on the dataset page: https://huggingface.co/datasets/espsluar/crawlerlm-html-to-json.texttext-generationn<1K2 likes27 downloads10mo agoHugging Face09domofon /50k-HTML-PRETRAIN 50k-HTML-PRETRAIN Pretrain-style pairs: an English site assignment and a complete HTML document that implements it. Each row is one assignment (user) and one HTML page (assistant). A handful of rows include a PNG screenshot of the page; the rest leave screenshot empty. Split split n train 57,617 10 rows have a PNG in screenshot / images/. The other rows have a null screenshot. from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/domofon/50k-HTML-PRETRAIN.imagetext-generation10K<n<100K0 likes27 downloads1mo agoHugging Face10AI-Culture-Commons /philosophy-culture-translations-html-csv AI-Culture Philosophy and Culture Translations CSV + HTML Corpus The corpus contains an exceptionally diverse range of cultural, philosophical, and literary texts, available in 12 major languages. Among other topics, there is extensive engagement with the ethics and aesthetics of artificial intelligence and its cultural and philosophical implications, as well as connections between AI and philosophy of language and philosophy of mind. This project is maintained by a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/philosophy-culture-translations-html-csv.imagetranslation1K<n<10K2 likes23 downloads1y agoHugging Face11sanjaymalladi /dataviz-charts-html-dataset DataViz Individual Charts HTML Dataset 120 individual chart HTML files — 12 chart types × 10 themes — for teaching models to generate themed single-chart visualizations. Structure data/train-00000-of-00001.parquet — Main dataset in Parquet format README.md — Dataset card Columns Column Type Description filename string treemap-tech-innovation.html title string e.g. "Bar Chart — Midnight Galaxy" chart_type string bar, line, area, pie… See the full description on the dataset page: https://huggingface.co/datasets/sanjaymalladi/dataviz-charts-html-dataset.tabulartext-generationn<1K0 likes18 downloads3mo agoHugging Face12tkattkat101 /html-compression HTML Compression Dataset This dataset is designed for fine-tuning language models to understand and generate compressed HTML notation. Dataset Description The dataset contains pairs of standard HTML and compressed HTML using a custom compression algorithm that: Abbreviates common HTML tags (div → d, span → s, etc.) Shortens attribute names (class → c, style → st, etc.) Compresses CSS patterns (display: flex → d:f, etc.) Removes unnecessary whitespace and comments… See the full description on the dataset page: https://huggingface.co/datasets/tkattkat101/html-compression.texttext-generation1K<n<10K0 likes16 downloads1y agoHugging Face13erv1n /geo_html_200 GEO HTML 200 Dataset A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research. Features Column Description doc_id Unique document identifier url Source URL cleaned_text Parsed plain text content cleaned_text_length Character count query Associated search query title Document title topic_tags Topic classification Usage from datasets import load_dataset ds = load_dataset("erv1n/geo_html_200") tabulartext-generationn<1K0 likes13 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.