datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-protein-binder-design
Claude protein binder design — data release v1.0
1,440 de novo miniprotein binders (50 to 120 residues) against 16 targets, designed by two Claude models operating as autonomous protein-design agents (Mythos Preview, 900 designs; Opus 4.8, 540 designs) and characterized at two contract research organizations, Adaptyv Bio (cell-free expression; SPR/BLI kinetics with the design immobilized) and Twist Bioscience (Fc-fusion expression; capture SPR with a six-point antigen… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design.GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.PubLayNet
Dataset Card for PubLayNet
Dataset Summary
PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions.
Supported Tasks and Leaderboards
The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.Design2CodeThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the dataset in the huggingface format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
Example Usage
For example, you… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Design2Code.DesignBenchResultsChip_Design
AURA-1: designing an edge-AI chip for a neckband headphone
Working files from a solo, from-scratch attempt to specify and prototype an
edge-AI inference SoC for wireless neckband headphones: a resident,
speech-native conversational model on the device, the cloud called as a tool for
facts. The novel block is the NPU; Wi-Fi, Bluetooth, codec, ANC and PMU are
sourced, not designed.
The author is a process engineer learning chip design. The archive is
deliberately complete: the… See the full description on the dataset page: https://huggingface.co/datasets/sangramrout/Chip_Design.CreativePSD
Dataset Card for CreativePSD
Dataset Summary
CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images.
This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.design_bench_dataPKU-PosterLayout
Dataset Card for PKU-PosterLayout
Dataset Summary
PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.Rico
Dataset Card for Rico
Dataset Summary
Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.design-patents-not-in-impact
US Design Patents Not Included in IMPACT (2008-2026)
Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are
absent from the AI4Patents/IMPACT dataset.
IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that
IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period,
plus 4,824 patents from years IMPACT does cover but did not include. There is no patent
overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.vargov-design-catalog
Vargov®Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov®Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.sell-designer-watches
Sell designer watches
A Canonical Explainer: Selling Your Designer Watch at King Gold & Pawn KG_SUNSET_PARK Introduction: Understanding the Value of Your Timepiece Designer watches are more than mere instruments for telling time; they are intricate works of art, engineering marvels, and often significant investments. Crafted by renowned horological houses, these timepieces carry prestige, heritage, and intrinsic value derived from their materials, movements, and brand legacy. When… See the full description on the dataset page: https://huggingface.co/datasets/CollateralAnalytics/sell-designer-watches.DesignBenchDesigned-Vocalizations-Dataset
Designed Vocalizations Dataset
Paper · Demo & audio samples
The Designed Vocalizations Dataset supports voice conversion for designed vocalizations
— monster growls, robotic voices, and other sound-designed timbres — an area left
underexplored by benchmarks that focus on natural human speech. It curates diverse raw vocal
sources (speech and animal / non-linguistic sounds) and applies professional vocal-effects
processing to produce corresponding effect-modified variants. A… See the full description on the dataset page: https://huggingface.co/datasets/NCSOFT/Designed-Vocalizations-Dataset.PrismLayersPro
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
We introduce PrismLayersPro, a 20K high-quality multi-layer transparent image dataset with rewritten style captions and human filtering.
PrismLayersPro is curated from our 200K dataset, PrismLayers, generated via MultiLayerFLUX.
Dataset Structure
📑 Dataset Splits (by Style)
The PrismLayersPro dataset is divided into 21 splits based on visual style categories.Each… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PrismLayersPro.DesignCoder
DesignCoder UI Bench 200 — API 模型对比
7 个 API 模型在 DesignCoder 200 题 UI 生成基准上的产物与评分。
每条记录包含:任务 prompt、模型生成的单文件 HTML、渲染截图,以及三族 rubric 的逐条判定。
评测日期 2026-09-26 · judge = deepseek-v4.1-flash-expires-on-0910 · 生成状态:完整:每个模型 200 题全部生成并评分。
分数
按各模型已完成题目平均(覆盖不同时不可横向比较):
模型
网关 id
已生成
Overall
Overall(no-VSD)
Prompt Fit
Frozen
Landing
Dashboard
GPT-5.2
api_azure_openai_gpt-5.2
200
89.90
88.38
88.33
86.37
92.09
85.83
DeepSeek-V4 Flash
deepseek-v4-flash
200
89.17
87.40… See the full description on the dataset page: https://huggingface.co/datasets/xingxm/DesignCoder.graphene-design-universe-256k
Graphene Design Universe 256K
256,000 unrelaxed atomistic graphene designs, with images, full periodic cells, standard extended XYZ coordinates, reproducible recipes, geometry-quality flags and a geometry-similarity explorer.
Interactive 256K explorer · Preserved 64K release · 4K movie
This expansion preserves all 64,000 prior designs, IDs and coordinate-file bytes and adds 192,000 new designs. It broadens the original 16 groups and adds eight hybrid motif groups. The… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/graphene-design-universe-256k.design_bench_dataNovel_NLRP3_Inhibitor_Designs
Novel NLRP3 Inhibitors — GA-II Designed Ligand–Receptor Complexes
Why this target matters. NLRP3 sits upstream of IL-1β in gout, cardiovascular, metabolic and neurodegenerative disease, yet after two decades of effort no small-molecule NLRP3 inhibitor has reached approval.
176 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the allosteric site of the NLRP3 inflammasome NACHT module.
Each molecule was constructed against this pocket… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/Novel_NLRP3_Inhibitor_Designs.Design2Code-HARDThis dataset consists of 80 extra difficult webpages from Github Pages, which challenges SoTA multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the "easy" version of the Design2Code testset here
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.fashion_designCGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.Magazine
Dataset Card for Magazine
Dataset Summary
Magazine is a magazine layout dataset released with Content-aware Generative Modeling of Graphic Design Layouts. The paper studies graphic layout generation conditioned on visual and textual content and introduces a large-scale magazine layout dataset with fine-grained layout annotations and keyword labels.
Supported Tasks and Leaderboards
The dataset supports content-aware layout generation, graphic… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Magazine.FSHR_Agonist_Designs_8I2G
FSHR Allosteric Agonist Designs — 8I2G (GA-II)
292 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the transmembrane allosteric pocket of the follicle-stimulating hormone receptor (FSHR), using PDB 8I2G as the receptor.
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. All 292 are profiled for… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/FSHR_Agonist_Designs_8I2G.SMARCA2_VHL_MolecularGlue_Designs_6HAY
SMARCA2–VHL Molecular Glue Designs (PDB 6HAY)
884 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the SMARCA2 bromodomain–VHL interface of the ternary complex 6HAY (2.24 Å), so that one molecule spans both partners.
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. Each is supplied as a… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/SMARCA2_VHL_MolecularGlue_Designs_6HAY.Design2Code-hfThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
See the dataset in the raw files format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
FKBP12_F36V_Binder_Designs_1BL4
FKBP12 F36V Binder Designs (PDB 1BL4)
165 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the engineered FKBP12 F36V bump-hole cavity of 1BL4 (1.9 Å).
Each molecule was constructed against this pocket rather than selected from a compound library — docking (AutoDock Vina) came afterwards, to place and score the generated molecules in the site. Each is supplied as a complete protein–ligand complex.
Molecules were generated by the… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/FKBP12_F36V_Binder_Designs_1BL4.fashion_design_qa
