Team Ai
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sfd-anonymous /html-table-reconstruction-benchmark HTML Table Reconstruction Benchmark This repository contains the 100-sample HTML table reconstruction benchmark artifacts used for the paper's SFD MMD vs. EdgarTools vs. to_markdown comparison. Each sample starts from a synthetic SEC-style table and evaluates whether a model can reconstruct faithful HTML from a parser-specific markdown representation. The uploaded artifacts are the saved benchmark outputs used for the reported table; no model calls were rerun during upload.… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/html-table-reconstruction-benchmark.table-question-answeringn<1K0 likes198 downloads5mo agoHugging Face02PatrickHaller /cc-html-tokens-qwen3 CommonCrawl HTML, tokenised Packed token ids for language-model pretraining on raw web markup (not extracted text). tokens 15.18B shards 77 tokenizer Qwen/Qwen3-0.6B dtype uint32, flat, no header document separator EOS token id preprocessing light HTML clean (scripts/styles stripped), capped at 32,768 tokens/page Shards are not uniform in size: 00000-00047 hold ~92M tokens each (8k pages per source WARC file), 00048 onward hold ~370M each (29k pages… See the full description on the dataset page: https://huggingface.co/datasets/PatrickHaller/cc-html-tokens-qwen3.texttext-generationn<1K0 likes138 downloads11d agoHugging Face03DogukanUrker /ui-distill-html-648 ui-distill-html-648 648 single-file HTML UI components, generated by Ornith-1.0-35B on a single RTX 3060 12GB, paired with the build request that produced each one. Built to fine-tune a 3B model into writing UI (DogukanUrker/ui-distill-3b), but it stands on its own — distill your own student from it. The interesting part The instructions in this dataset are not the prompts that generated the HTML. The teacher was driven by a 678-character system prompt: use the… See the full description on the dataset page: https://huggingface.co/datasets/DogukanUrker/ui-distill-html-648.texttext-generationn<1K2 likes88 downloads3mo agoHugging Face04thekevinscott /geocities-prompt-html GeoCities prompt → HTML — fine-tune Fine-tunes Gemma-4-E2B-it (LoRA) to generate a full, vintage-style HTML page from a plain-language description. This repo holds the dataset and the training scripts so the whole thing runs from one place. What's in here dataset.jsonl — the training data: one {"prompt": ..., "html": ...} per line. train_geocities.py — training entrypoint (loads this JSONL format). train-geocities-5090.sh — launch tuned for a 32 GB card (bf16… See the full description on the dataset page: https://huggingface.co/datasets/thekevinscott/geocities-prompt-html.texttext-generation1K<n<10K2 likes50 downloads4mo agoHugging Face05Boakpe /deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript. texttext-generationn<1K1 likes32 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.