datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DesignCoder
DesignCoder UI Bench 200 — API 模型对比
7 个 API 模型在 DesignCoder 200 题 UI 生成基准上的产物与评分。
每条记录包含:任务 prompt、模型生成的单文件 HTML、渲染截图,以及三族 rubric 的逐条判定。
评测日期 2026-09-26 · judge = deepseek-v4.1-flash-expires-on-0910 · 生成状态:完整:每个模型 200 题全部生成并评分。
分数
按各模型已完成题目平均(覆盖不同时不可横向比较):
模型
网关 id
已生成
Overall
Overall(no-VSD)
Prompt Fit
Frozen
Landing
Dashboard
GPT-5.2
api_azure_openai_gpt-5.2
200
89.90
88.38
88.33
86.37
92.09
85.83
DeepSeek-V4 Flash
deepseek-v4-flash
200
89.17
87.40… See the full description on the dataset page: https://huggingface.co/datasets/xingxm/DesignCoder.analog-design-benchAnalog Design Bench
Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks
Website
·
GitHub
Analog Design Bench contains 50 analog and mixed-signal circuit design tasks in
the SkyWater SKY130 process. In each task an agent designs a transistor-level
circuit and delivers it as a SPICE netlist, /app/circuit.spi. A separate
verifier, which the agent never sees, simulates the netlist with ngspice and
checks it against every specification; a task is solved… See the full description on the dataset page: https://huggingface.co/datasets/Arcadia-2026/analog-design-bench.TTS-Voice-Design-Benchmark
TTS Voice Design Benchmark
🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Design is a high-quality benchmark of 1,000 character voice-design
tasks spanning a broad range of media genres and real-world creative use
cases. It evaluates whether a text-to-speech model can turn an open-ended
character profile into a distinctive, appropriate, and usable voice.
Unlike benchmarks built around a fixed set of speakers or isolated acoustic
attributes, this dataset covers complete… See the full description on the dataset page: https://huggingface.co/datasets/BreezeBlue/TTS-Voice-Design-Benchmark.mtg-oracle-design-descriptions-v2
MTG Oracle descriptions — small holdout v2
Training: 316,372 descriptions / 35,744 cards. Final evaluation: 400 descriptions from 400 disjoint rules families. All 3,285 other descriptions of held-out families are preserved in heldout-reserve.jsonl, excluded from training and headline evaluation. No validation split.
Same reviewed descriptions, editable JSON/Oracle targets, names, rarity, loyalty and provenance as v1; only assignment changed. Choose messages, sft… See the full description on the dataset page: https://huggingface.co/datasets/vishvananda/mtg-oracle-design-descriptions-v2.mtg-oracle-design-descriptions
MTG Oracle descriptions — oracle-full-v1
Status: complete. Train: 250,587; validation: 31,852; test: 37,618.
Natural-language card requests paired with editable card JSON and Oracle text. Targets include names and rarity. Unspecified characteristics are editable suggestions, not uniquely correct answers.
Use messages for ordinary chat SFT, sft for prompt/completion masking, or records for source IDs, detail levels, specified/proposed fields and review provenance. The local… See the full description on the dataset page: https://huggingface.co/datasets/vishvananda/mtg-oracle-design-descriptions.game-design-pattern-core-collection
Game Design Patterns Dataset
Original Source Attribution
This dataset is derived from the work of Staffan Björk and Jussi Holopainen. The original content comes from "Patterns in Game Design," published by Charles River Media in 2005.
Original authors: Jussi Kuittinen, Staffan Björk and Jussi Holopainen
Original format: HTML documents publicly available at: https://www.researchgate.net/publication/379683418_collection.zip
Citation: Bjork, S., & Holopainen, J. (2005).… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/game-design-pattern-core-collection.qi-design-submissions
QI Design Submissions
中文
Submit one JSON file per pull request under submissions/<unique-name>.json. Download the template from the benchmark's Submit tab. A file must have exactly schema_version: "qi-design-v1", method and inputs.
inputs contains n_field_periods (2, 3 or 4), r_cos and z_sin. Each array contains one or more matrices of shape 5x9, with matching counts. All designs share NFP. Finite numbers only; r_cos[0][0:4] and z_sin[0][0:5] must be zero in each matrix. The… See the full description on the dataset page: https://huggingface.co/datasets/SII-AI4Fusion/qi-design-submissions.am-session-sharing-design
Agent Manager session — designing session sharing
Access: public. Anyone with the link can read this trace. Public is the only mode
where the Hub's trace viewer works for every visitor regardless of account tier.
A single real Agent Manager
session, exported as a shareable trace. This is the working session in which the
session-sharing design itself was researched and written — surveying how five coding-agent
harnesses persist sessions, testing what the Hub's trace viewer can… See the full description on the dataset page: https://huggingface.co/datasets/thomwolf/am-session-sharing-design.am-session-sharing-design-noimg
Agent Manager session — designing session sharing
Access: public. Anyone with the link can read this trace. Public is the only mode
where the Hub's trace viewer works for every visitor regardless of account tier.
A single real Agent Manager
session, exported as a shareable trace. This is the working session in which the
session-sharing design itself was researched and written — surveying how five coding-agent
harnesses persist sessions, testing what the Hub's trace viewer can… See the full description on the dataset page: https://huggingface.co/datasets/thomwolf/am-session-sharing-design-noimg.nonattainment-county-designations-by-pollutant
Current EPA nonattainment county designations by pollutant
Canonical, always-current version: https://referencesource.org/nonattainment-county-designations-by-pollutant/
Machine-readable: https://referencesource.org/nonattainment-county-designations-by-pollutant/data.json — this mirror is a point-in-time copy.
Last verified: 2026-10-05
Stale after: 2027-04-03 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 463
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nonattainment-county-designations-by-pollutant.DesignAsCode-training-data
DesignAsCode Training Data
Training data for the DesignAsCode Semantic Planner.
Overview
Samples
19,479
Format
JSONL (one JSON object per line)
Size
~145 MB
Data Source
Each sample corresponds to a real graphic design from the Crello dataset. We distilled structured design semantics from each original design using GPT-4o and GPT-o3, taking the original design, its individual layers, and design metadata as input.
The distillation produces:… See the full description on the dataset page: https://huggingface.co/datasets/Tony1109/DesignAsCode-training-data.Bio-Design-ProcessThis dataset works even though it may not be the cleanest in regards to organization. I'm working on cleaning it up for better performance, but it should still work as long as you don't overtrain on it.
ui-design-audit-dataset
UI Design Audit Screenshot Benchmark v2.1
A reproducible synthetic benchmark of 3,000 mobile and web UI screenshots labeled across 12 design-risk categories.
Splits
train: 2,400
validation: 300
test: 300
Labels
small_touch_targets
low_contrast
action_overload
navigation_overload
form_friction
content_density
responsive_risk
modal_overuse
deep_scrolling
weak_hierarchy
interaction_overload
mobile_web_mismatch
Data creation
Every… See the full description on the dataset page: https://huggingface.co/datasets/newazhala/ui-design-audit-dataset.agents-md-agent
agents-md-agent — "Hub Scout"
An agent whose entire instruction set lives in one human-readable file,
AGENTS.md. The code (app.py) is just the wiring: it loads AGENTS.md
verbatim as the system prompt and gives the model two tools that hit the
real Hugging Face Hub. Edit AGENTS.md and the agent's behavior changes —
no code changes. This is the "author markdown, push, it's live" pattern.
Run locally
git clone… See the full description on the dataset page: https://huggingface.co/datasets/designfailure/agents-md-agent.M-DESIGN-Knowledge-Base
M-DESIGN Knowledge Base and Model Artifacts
This dataset contains the released SQLite model-performance databases and model
artifacts used by M-DESIGN, the method from "Beyond Model Base Retrieval:
Weaving Knowledge to Master Fine-grained Neural Network Design".
Contents
Each .db file has a model_records table. The first six columns encode
fine-grained neural design choices and the final two columns store the measured
score and standard deviation. Each task/dataset… See the full description on the dataset page: https://huggingface.co/datasets/jilwang804/M-DESIGN-Knowledge-Base.Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.arch-design-sft
arch-design-sft: verified architecture-design SFT data
Supervised fine-tuning data for neural architecture design treated as structured graph editing. Each row pairs a natural-language design spec and a serialized starting graph with a reference action plan, and every row is re-graded by a deterministic verifier before it is written: structural blockers, parameter budgets and bands, required layer families. Ten task families across six design-from-spec and four edit-in-place… See the full description on the dataset page: https://huggingface.co/datasets/neurarch-ai/arch-design-sft.designer-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.ParaSFT-designer
ParaSFT Designer
English | 中文
Overview
ParaSFT Designer is a private supervised fine-tuning dataset for ParadoxGPT-Designer-4B, the ParadoxGPT specialist model for experiment design, evidence planning, and claim-to-experiment mapping.
Designer annotation pipeline over ParaPaper context packs, covering claim-to-evidence mapping, experiment argument planning, ablation design, sufficiency critique, and interpretation boundaries.
Each example is an instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/bhxdianzhang/ParaSFT-designer.ptdbench-reward-design-task-grpo-broadcast-dataset
PTDBench dataset snapshot: task_grpo_broadcast
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; bytes identical to task_function_call.
License: Apache-2.0
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-task-grpo-broadcast-dataset.2026-08-26-sonnet45-post-action-retrospection-natural-turn-design
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
field
value
experiment
synth post_action_retrospection run — per-stage snapshots (resumable generation cache)
date_generated
20260826_152715
constitution
constitutions/claude_distilled_12_principles_mid/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8
models
per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.design-3d-003
Design 3D-003 — Production Dataset
Production synthetic dataset for training a from-scratch Transformer
to generate structured 3D architectural scenes for Blender.
Dataset
Total scenes: 20,000
Train: 14,000
Validation: 3,000
Test: 3,000
Unique scene hashes: 20,000
Schema version: 1.1.0
Generator: design_3_003_production_generator_v1
Production counts
Objects: 185,014
Openings: 65,334
Relations: 986,163
Files
train.json — 14,000… See the full description on the dataset page: https://huggingface.co/datasets/quachtohai/design-3d-003.web-design-diamond
Web Design Diamond — dataset SFT
Dataset para entrenar LLMs (chicos: 2B-8B) que, dado un pedido en lenguaje natural, generan
UNA pagina index.html autocontenida (Tailwind CSS por CDN + JS vanilla embebido) que funciona
y se ve premium, rapido. Texto -> codigo (no imagen -> codigo).
Que tiene de distinto
Razonamiento (thinking) secuencial y sin leakage: el modelo razona que va a hacer ANTES
de implementar; el thinking previo a una tool no menciona tokens… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/web-design-diamond.MF-Design
Data for MF-Design
This repository hosts the datasets for the paper "Repurposing AlphaFold3-like Protein Folding Models for Antibody Sequence and Structure Co-design" (MF-Design).
The main code repository can be found at MF-Design.
Directory Structure
The ./data/ directory contains all the necessary data for running and evaluating the models.
Raw and Processed Data: Includes the original data used for training in ./data/raw_data.tar.zst and its processed versions in… See the full description on the dataset page: https://huggingface.co/datasets/clorf6/MF-Design.design_compiler_pdf_datasetfresh-by-design-cosmetic-lexicon
Fresh by Design Cosmetic Lexicon v1.0
This dataset is a small bilingual lexicon of cosmetic freshness terms used to describe Sostenica's "Fresh by Design" / "Fresco Por Diseño" formulation philosophy. Version 1.0 is aligned with Sostenica's public Glosario de Formulación.
The dataset is intentionally small. Its purpose is not model training at scale. Its purpose is structured terminology: a machine-readable reference for a freshness-oriented cosmetic vocabulary around formulation… See the full description on the dataset page: https://huggingface.co/datasets/sostenica/fresh-by-design-cosmetic-lexicon.design_compiler_md_datasetsystem_design
data ingestion product Nexus
Description
It's about the project Nexus whose task is to gather the data from all the media, video, digital channels of a company and keep ingesting it thus keeping the knowledge base embeddings up to date.
Format
This dataset is in alpaca format.
Creation Method
This dataset was created using the Easy Dataset tool.
Easy Dataset is a specialized application designed to streamline the creation of fine-tuning datasets for… See the full description on the dataset page: https://huggingface.co/datasets/SaffalPoosh/system_design.bit_new_design_1Creative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset
Creative Design Studio Interaction Body Language Recognition Video Dataset
In today's creative design industry, understanding the complex non-verbal communication among team members is crucial. However, existing body language recognition methods perform limitedly in dense interactive environments, struggling to accurately interpret subtle body movements and postures. This video dataset aims to tackle the technical challenges of body language analysis in creative discussions… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Creative-Design-Studio-Interaction-Body-Language-Recognition-Video-Dataset.
