Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes2.2k downloads4y agoHugging Face02masoudjs /c4-en-html-with-metadata-ppl-cleanFile list: "c4-en-html_cc-main-2019-18_pq00-000.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-001.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-002.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-003.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-004.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-005.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-006.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-007.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-008.jsonl.gz", "c4-en-html_cc-main-2019-18_pq00-009.jsonl.gz"… See the full description on the dataset page: https://huggingface.co/datasets/masoudjs/c4-en-html-with-metadata-ppl-clean.tabular10K<n<100K1 likes560 downloads4y agoHugging Face03SaulLu /Natural_Questions_HTMLThis is a dataset extracted from the Natural Questions dataset This dataset is currently under development text10K<n<100K0 likes415 downloads5y agoHugging Face04common-pile /stackv2_html_filteredtext1M<n<10M3 likes316 downloads1y agoHugging Face05zstanjj /HtmlRAG-train Dataset Information We release the training data used in HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieval Results in RAG Systems. The training data is sampled from 5 widely used QA datasets, including ASQA, Hotpot-QA, NQ, Trivia-QA, and MuSiQue. You can access the original dataset here. We apply Bing search API in the US-EN region to search for relevant web pages, and then we scrap static HTML documents through URLs in returned search results. We provide the URLs and… See the full description on the dataset page: https://huggingface.co/datasets/zstanjj/HtmlRAG-train.textquestion-answering10K<n<100K1 likes294 downloads1y agoHugging Face06SaulLu /Natural_Questions_HTML_Toytextn<1K0 likes272 downloads5y agoHugging Face07SaulLu /Natural_Questions_HTML_reduced_alltext10K<n<100K4 likes240 downloads5y agoHugging Face08sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes198 downloads5mo agoHugging Face09bs-modeling-metadata /c4-en-html-with-metadatatabular10M<n<100M14 likes153 downloads4y agoHugging Face10PatrickHaller /cc-html-tokens-qwen3 CommonCrawl HTML, tokenised Packed token ids for language-model pretraining on raw web markup (not extracted text). tokens 15.18B shards 77 tokenizer Qwen/Qwen3-0.6B dtype uint32, flat, no header document separator EOS token id preprocessing light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.texttext-generationn<1K0 likes138 downloads11d agoHugging Face11Reubencf /frontend-html-tailwind-js This dataset is a remastered version of Reubencf/frontend-coding prepared using Adaption's Adaptive Data platform. frontend_html_tailwind_js This dataset contains pairs of user prompts and generated frontend code solutions using HTML, Tailwind CSS, and JavaScript. The samples cover a wide range of web development tasks, including landing pages, portfolios, ecommerce sites, and interactive components with smooth scrolling and animations. Each entry demonstrates practical… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/frontend-html-tailwind-js.textn<1K1 likes105 downloads6mo agoHugging Face12DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K2 likes88 downloads3mo agoHugging Face13theprint /MultiRoundConvos-Code-JS-HTML-CSS-Pythontext1K<n<10K0 likes72 downloads10mo agoHugging Face14franktheglock /html-to-markdown-10000-pairs HTML to Markdown 10,000 Pairs This dataset contains 10,000 synthetic paired examples for HTML-to-Markdown conversion. Structure pairs/train.jsonl: 9,000 examples pairs/validation.jsonl: 500 examples pairs/test.jsonl: 500 examples manifest.json: generation manifest and split counts Each JSONL row includes: id: example identifier html: source HTML string markdown: target Markdown string metadata: per-example metadata Example {"id":"sample-0000000"… See the full description on the dataset page: https://huggingface.co/datasets/franktheglock/html-to-markdown-10000-pairs.texttranslation10K<n<100K0 likes51 downloads7mo agoHugging Face15thekevinscott /geocities-prompt-html GeoCities prompt → HTML — fine-tune Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page from a plain-language description. This repo holds the dataset and the training scripts so the whole thing runs from one place. What's in here dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line. train_geocities.py — training entrypoint (loads this JSONL format). train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.texttext-generation1K<n<10K2 likes50 downloads4mo agoHugging Face16ttbui /alpaca_webgen_htmltextn<1K1 likes42 downloads3y agoHugging Face17Jiraya /html_to_json_information_extraction_dataset HTML to JSON Information Extraction Dataset Description The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON. These HTML have been sourced (scraped) from about 25 companies' career pages. The dataset contains three splits - train, test, unseen_test. This dataset has been built to fine tune SLMs & LLMs for the information extraction task. train split This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.text1K<n<10K2 likes35 downloads2y agoHugging Face18Boakpe /deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript. texttext-generationn<1K1 likes32 downloads9mo agoHugging Face19Juliankrg /HTML_CSS_CodeDataSet_100ktext100K<n<1M5 likes30 downloads2y agoHugging Face20TeichAI /gemini-3-flash-preview-standalone-html-1kThis dataset was made by querying Gemini 3 Flash Preview to train models on web development for games, websites, and web apps in html + inline css and javascript. Prompt breakdown: First 800 targeted at standalone html websites, games, and apps. The last 200 prompts are focused on science (chemistry, physics, biology), math, philosophy, multilingual tasks, and complex reasoning tasks. All prompts were generated by gemini 3 pro preview Stats: Cost: $ 4.73 (USD) Total tokens (input +… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/gemini-3-flash-preview-standalone-html-1k.text1K<n<10K2 likes29 downloads10mo agoHugging Face21BarBar288 /html 📚 HTML Tags Complete Reference Dataset A comprehensive, structured dataset of HTML tags with detailed attributes, examples, and metadata. Perfect for learning HTML, building educational tools, or training ML models on web development concepts. 📊 Dataset Description This dataset contains 85+ HTML tags with rich metadata including tag purpose, category, attributes, usage examples, and more. Each entry provides complete information about an HTML tag, making it an ideal… See the full description on the dataset page: https://huggingface.co/datasets/BarBar288/html.texttext-classificationn<1K0 likes21 downloads8mo agoHugging Face22ttbui /html_alpacatextn<1K7 likes19 downloads3y agoHugging Face23hiimivantang /chunked-enwiki-ns0-20250301-enterprise-html20250301.jsonl contains preprocessed text chunks from the wikipedia 20250301 html dump. dead_letter_queue.jsonl contains articles that couldn't be processed due to formatitng or missing tags. en_redirection_map.pkl can be deserialize to avoid building redirection_map everytime preprocessing_wikipedia_html_dump runs. tabular10M<n<100M0 likes19 downloads2y agoHugging Face24ttbui /alpaca_data_with_html_outputtext10K<n<100K7 likes18 downloads3y agoHugging Face25jawerty /html_datasettextn<1K9 likes16 downloads4y agoHugging Face26jasong03 /vov_thegioi_htmltext100K<n<1M0 likes13 downloads2y agoHugging Face274cast /small-fable-html-dataset Dataset Card for 4cast/small-fable-html-dataset This dataset provides a few high-quality examples of non-thinking outputs by Claude Fable 5. They have been taken from various sources such as ultimateplay.com. Dataset Details Dataset Sources Repository: ultimateplay.com, claude.ai, arena.ai textn<1K0 likes11 downloads4mo agoHugging Face28jasong03 /vov_suckhoe_htmltext10K<n<100K0 likes9 downloads2y agoHugging Face29guardiancc /image-to-htmltextn<1K6 likes8 downloads3y agoHugging Face30jawerty /htmlembedtextn<1K0 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.