datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GenPoster100K
Dataset Card for GenPoster100K
Dataset Summary
GenPoster-100K is a large-scale dataset for content-aware graphic layout generation introduced in the SEGA paper.
The paper describes it as a high-quality poster dataset with layer-parseable source materials and rich metadata.
This repository provides a Hugging Face datasets loader implementation that reads the source release (BruceW91/GenPoster-100K) and exposes normalized examples with:
poster background image… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/GenPoster100K.graphwalks
GraphWalks: a multi hop reasoning long context benchmark
In Graphwalks, the model is given a graph represented by its edge list and asked to perform an operation.
Example prompt:
You will be given a graph as a list of directed edges. All nodes are at least degree 1.
You will also get a description of an operation to perform on the graph.
Your job is to execute the operation on the graph and return the set of nodes that the operation results in.
If asked for a breadth-first… See the full description on the dataset page: https://huggingface.co/datasets/openai/graphwalks.PubLayNet
Dataset Card for PubLayNet
Dataset Summary
PubLayNet is a large document layout analysis dataset built by automatically matching XML representations and PDF content from more than one million PubMed Central Open Access articles. It contains more than 360,000 document images with COCO-style annotations for common layout elements such as text, title, list, table, and figure regions.
Supported Tasks and Leaderboards
The dataset supports document… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PubLayNet.PKU-PosterLayout
Dataset Card for PKU-PosterLayout
Dataset Summary
PKU-PosterLayout is a content-aware visual-textual poster layout benchmark released with PosterLayout: A New Benchmark and Approach for Content-aware Visual-Textual Presentation Layout. The paper defines the task as arranging predefined text, logo, and underlay elements on a non-empty poster canvas while considering both inter-element and inter-layer relationships. The original benchmark contains 9,974… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PKU-PosterLayout.CreativePSD
Dataset Card for CreativePSD
Dataset Summary
CreativePSD is the PSD-derived graphic design dataset released with PSDesigner. Each example is a poster archive containing PSD tree text, structured layer metadata, tool-call trajectories, source image resources, and stepwise rendered images.
This loader keeps the contents of each poster_*.zip archive: all metadata text/JSON files, all raw_resource images, all rendering_imgs images, and a manifest of every member in… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CreativePSD.Mintaka_Graph_Features_T5-xl-ssm
Dataset Card for "Mintaka_Graph_Features_T5-xl-ssm"
More Information needed
graphene-design-universe-256k
Graphene Design Universe 256K
256,000 unrelaxed atomistic graphene designs, with images, full periodic cells, standard extended XYZ coordinates, reproducible recipes, geometry-quality flags and a geometry-similarity explorer.
Interactive 256K explorer · Preserved 64K release · 4K movie
This expansion preserves all 64,000 prior designs, IDs and coordinate-file bytes and adds 192,000 new designs. It broadens the original 16 groups and adds eight hybrid motif groups. The… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/graphene-design-universe-256k.Rico
Dataset Card for Rico
Dataset Summary
Rico is a mobile app UI dataset for building data-driven design applications. The original dataset mines Android apps at runtime and exposes visual, textual, structural, and interactive design properties from more than 9.3k apps across 27 categories and more than 66k unique UI screens. This packaging provides metadata, screenshots, view hierarchies, and semantic annotations as separate configs.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/Rico.Graph-AlgorithmsText-Attributed-Graphs
Overview
This dataset covers the encoder embeddings and prediction results of LLMs of paper 'Model Generalization on Text Attribute Graphs: Principles with Lagre Language Models', Haoyu Wang, Shikun Liu, Rongzhe Wei, Pan Li.
Dataset Description
The dataset structure should be organized as follows:
/dataset/
│── [dataset_name]/
│ │── processed_data.pt # Contains labels and graph information
│ │── [encoder]_x.pt # Features extracted by different encoders
│… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/Text-Attributed-Graphs.dach_bike_graph
DACH Bike + Rail Routing Graph
A prebuilt, ready-to-route cycling + railway graph covering Germany, Austria, and Switzerland (DACH), stored as lat/lon-tiled GeoParquet. Built from OpenStreetMap by the Bike Route Optimizer for flat-preferring, surface-aware bike routing that can also hop on a train uphill.
Everything routing needs is baked in — node elevations and full 3D edge geometry — so an application downloads this once and routes offline, with no Overpass and no elevation… See the full description on the dataset page: https://huggingface.co/datasets/MichaelMedek/dach_bike_graph.GraphRAG-Bench
GraphRAG-Bench : A Comprehensive Benchmark for Evaluating Graph Retrieval-Augmented Generation Models
🎉News •
📖About •
🏆Leaderboards •
🧩Task Examples
🔧Getting Started •
📬Contact •
📝Citation
This repository is for the GraphRAG-Bench project, a comprehensive benchmark for evaluating Graph Retrieval-Augmented Generation models.
🎉 News
[2025-05-25] We release GraphRAG-Bench, the benchmark for evaluating GraphRAG… See the full description on the dataset page: https://huggingface.co/datasets/GraphRAG-Bench/GraphRAG-Bench.szl-estate-graph
SZL Estate Graph v1
This is a deterministic publication bundle for SZLHOLDINGS/szl-estate-graph.
It turns the two receipted SZL Constellation topology documents into a typed,
trinity-connected graph suitable for graph-learning and drift comparison.
Publication target: SZLHOLDINGS/szl-estate-graph. Hub publication and its
immutable revision are provider evidence separate from this content bundle;
neither publication nor download implies model training, admission, or runtime
use.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/szl-estate-graph.GBC10M
Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs
GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs.
** The associated data point can be found at demo/water_tower.json
Description and data format
The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.GraphResidual-data
GraphResidual data package
Public research dataset package for the GraphResidual repository. It contains
source datasets, derived benchmark tables, processed graph tables, and graph
artifacts used by the ESM-2 and AMPLIFY-120M experiments.
Layout
data_raw/: source data and model assets used by the released experiments.
data_benchmark/: benchmark splits, candidate sets, and audit tables.
data_processed/: processed node, edge, and label tables.… See the full description on the dataset page: https://huggingface.co/datasets/Marcochris/GraphResidual-data.python-image-copilot-training-using-import-knowledge-graphs
Python Copilot Image Training using Import Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 216642
Size: 211.2 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-import-knowledge-graphs.llm-graph-poisoning-data
Generation-Time Poisoning of LLM-Generated Social Networks
This dataset contains synthetic personas, LLM-generated social graphs, cached
text embeddings, and evaluation metrics for clean generation and three
generation-time attack families. All names and profiles are synthetic and do
not represent real people.
Dataset variants
Variant
Nodes
Generator
Graph seeds per condition
Attack rates
p50
50
Qwen3-Max
10
10%, 20%, 30%, 40%, 50%
p200
200… See the full description on the dataset page: https://huggingface.co/datasets/Kevynf/llm-graph-poisoning-data.AnonyRAG
AnnoyRAG Dataset
The AnnoyRAG dataset, introduced in Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning, employs entity anonymization to isolate LLMs' parametric knowledge. This design enables more precise evaluation of how effectively LLMs integrate retrieved information in RAG systems.
Dataset Details
Dataset Description
The basic statistical information of the dataset is as follows:
Question Type
Difficulty Level… See the full description on the dataset page: https://huggingface.co/datasets/Youtu-Graph/AnonyRAG.kaggle-knowledge-graph
Kaggle Knowledge Graph
A knowledge graph built from Kaggle's public
Meta Kaggle (and Meta Kaggle Code
datasets). It links competitions, teams, submissions, users, notebooks,
datasets, discussion forums, tags, organizations, and notebook code invocations.
See the project repository for build scripts and
documentation on the graph schema, data model, and usage examples.
Release
Field
Value
Version
2026-09-30
Meta Kaggle snapshot
2026-09-30
Build code… See the full description on the dataset page: https://huggingface.co/datasets/habedi/kaggle-knowledge-graph.python-image-copilot-training-using-class-knowledge-graphs
Python Copilot Image Training using Class Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each row contains a png file in the dbytes column.
Rows: 312277
Size: 304.3 GB
Data type: png
Format: Knowledge graph using NetworkX with alpaca text box
Schema
The png is in the dbytes column:
{
"dbytes": "binary"… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-image-copilot-training-using-class-knowledge-graphs.PrismLayersPro
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
We introduce PrismLayersPro, a 20K high-quality multi-layer transparent image dataset with rewritten style captions and human filtering.
PrismLayersPro is curated from our 200K dataset, PrismLayers, generated via MultiLayerFLUX.
Dataset Structure
📑 Dataset Splits (by Style)
The PrismLayersPro dataset is divided into 21 splits based on visual style categories.Each… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/PrismLayersPro.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.mm-graph
Multimodal Graph Benchmark
Paper: https://huggingface.co/papers/2406.16321
Project Page: https://mm-graph-benchmark.github.io/
Code: https://github.com/mm-graph-benchmark/mm-graph-benchmark
This repo contains all the datasets used in "Multimodal Graph Benchmark".
THE DATASET IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE… See the full description on the dataset page: https://huggingface.co/datasets/mm-graph-org/mm-graph.rest-graph-searchCGL-Dataset-v2
Dataset Card for CGL-Dataset v2
Dataset Summary
CGL-Dataset v2 is an advertising-poster layout dataset released with Relation-Aware Diffusion Model for Controllable Poster Layout Generation. The paper argues that poster layouts should account for both visual-textual relationships and geometry relationships between elements. This version extends CGL-Dataset with richer element annotations, text annotations, and text features for controllable poster layout… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset-v2.graph-pannuke
Graph-PanNuke: A Cell-Graph Dataset for Nucleus Classification from PanNuke
Graph-PanNuke is a node-level classification dataset derived from the PanNuke pan-cancer histology dataset. We use all slides at 40× magnification. Each tissue patch is converted into a cell-graph where nodes represent detected cell nuclei and edges encode spatial proximity. The task is predicting the cell type of each nucleus across 5 classes. Note that node features describe cell morphology, texture… See the full description on the dataset page: https://huggingface.co/datasets/ogutsevda/graph-pannuke.sketch-graph-vitruvion
pretty_name: SketchGraphs, Vitruvion selection (rebuilt)
license: other
license_name: onshape-terms-of-use
license_link: https://www.onshape.com/legal/terms-of-use#your_content
size_categories:
- 1M<n<10M
tags:
- cad
- parametric-cad
- sketches
- geometric-constraints
- sketchgraphs
- vitruvion
SketchGraphs, Vitruvion selection (sg_filtered_unique.npy, rebuilt)
This is the dataset of Vitruvion (Seff et al., Vitruvion: A Generative Model of… See the full description on the dataset page: https://huggingface.co/datasets/benikm91/sketch-graph-vitruvion.CGL-Dataset
Dataset Card for CGL-Dataset
Dataset Summary
CGL-Dataset is a poster layout dataset released with Composition-aware Graphic Layout GAN for Visual-Textual Presentation Designs. The paper studies layout generation for a given image, emphasizing that both global semantics and spatial image composition affect where graphic elements should be placed. The original dataset contains 60,548 advertising posters with annotated layout information.
Supported… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/CGL-Dataset.math-graph
Math-Graph
Math-Graph is the dataset behind TheoremGraph, a unified, statement-level dependency
graph spanning both informal and formal mathematics. On the informal side it parses millions of
theorem-like environments from mathematics arXiv and recovers directed dependency edges within and
across papers; on the formal side it releases LeanGraph, an elaborator-level extraction of typed
declaration dependencies across 25 Lean 4 projects. The two graphs are bridged into one… See the full description on the dataset page: https://huggingface.co/datasets/uw-math-ai/math-graph.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.
