datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic_data_v5_finegrain_layout_relight_with_our_synthetic_data_coco_l_full_500kSQuADDS_Layouts
SQuADDS Layouts - versioned GDS artifacts for superconducting quantum hardware
SQuADDS Layouts is the geometry-artifact companion to
SQuADDS_DB, the
Superconducting Qubit And Device Design and Simulation Database. It provides
checksum-verified GDS files, stable geometry identities, and machine-readable
geometry metadata so a simulation result can be traced to the exact layout
that produced it.
Homepage: https://lfl-lab.github.io/SQuADDS/
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layouts.blender-3d-layout-v2-720p
Blender 3D Layout — 720p comparison groups (release v3.1)
Synthetic multi-view renders of objects placed into real Blender scenes, built so
that a change to the layout can be isolated: every camera pose is rendered
for the base layout, with one object moved, with an extra object added, and with
one object rotated in place — with the same camera matrix each time. Each
layout also comes with a box-orientation prompt in the style of BoxCtrl: every
object's 3D box drawn on a white… See the full description on the dataset page: https://huggingface.co/datasets/3d-layout/blender-3d-layout-v2-720p.funsd-layoutlmv3LayoutSAM
LayoutSAM Dataset
Overview
The LayoutSAM dataset is a large-scale layout dataset derived from the SAM dataset, containing 2.7 million image-text pairs and 10.7 million entities. Each entity is annotated with a spatial position (i.e., bounding box) and a textual description.
Traditional layout datasets often exhibit a closed-set and coarse-grained nature, which may limit the model's ability to generate complex attributes such as color, shape, and texture.… See the full description on the dataset page: https://huggingface.co/datasets/HuiZhang0812/LayoutSAM.synthetic-bedroom-layouts
Synthetic Bedroom Layouts
4,000 procedurally composed bedroom layouts with 27,025 furniture placements,
in the box format used by indoor scene-synthesis models (ATISS-style boxes.npz).
Furniture is drawn from Amazon Berkeley Objects (CC BY 4.0), so the whole
dataset is redistributable and usable commercially.
Nothing here derives from 3D-FRONT, 3D-FUTURE, or any dataset that restricts
redistribution. Where the numbers came from is written out, value by value, in
PROVENANCE.md.… See the full description on the dataset page: https://huggingface.co/datasets/Spatial1ntelligence/synthetic-bedroom-layouts.Orchestra_Layout_Dataset
Orchestral Score Layout Dataset
Ground-truth annotations for staff layout detection in orchestral music scores, covering two Tchaikovsky symphonies.
Contents
dataset_layout/
├── Tchai_4.pdf # Source PDF — Tchaikovsky Symphony No. 4
├── Tchai_4/ # Page images (PNG, one per page)
├── Tchai_4_csv_gt/ # Per-page CSV ground truth for Tchai_4
│
├── Tchai_6.pdf # Source PDF — Tchaikovsky Symphony No. 6
├── Tchai_6/… See the full description on the dataset page: https://huggingface.co/datasets/BowenC/Orchestra_Layout_Dataset.ndl-layout-dataset
NDL-DocL Kotenseki Layout Dataset (YOLO format)
A YOLO-formatted conversion of the kotenseki (pre-modern Japanese materials, 古典籍資料)
subset of the NDL-DocL dataset published by the National Diet Library of Japan (NDL).
Source dataset: https://github.com/ndl-lab/layout-dataset
Source images: NDL Digital Collections https://dl.ndl.go.jp/
国立国会図書館が公開する NDL-DocL データセットのうち、古典籍資料を
YOLO 形式(Ultralytics 互換)に変換したものです。
This is a modified/derived version. The bounding boxes were converted… See the full description on the dataset page: https://huggingface.co/datasets/nakamura196/ndl-layout-dataset.regulated_layout_dataset_v9_20260802RealDoc-Bench-Layout
RealDocBench-Layout
A 1,500-page document-layout benchmark for evaluating layout-detection
models on real-world documents. COCO-style annotations across 9 block
classes.
Contents
images/ — 1,500 page images (PNG / JPG / occasional WebP-as-PNG; see Caveats).
annotations/<pageId>.json — per-page COCO files, each with a single
image record, an annotations list, a categories list, and a
page_info block.
manifest.csv — pageId → domain + source URLs. The canonical row… See the full description on the dataset page: https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout.room-layout-planning-curated-v1
Room Layout Planning — curated pilot v1
39 个逐条检查并编写需求的房间布局任务,供实验流程验证与人工抽查。所有最终设计需求均为 AI 编写;没有人工标注或人工复核声明。
Split
条数
独立房屋
几何来源
train
26
26
InstructScene / 3D-FRONT
val_seen
7
7
InstructScene / 3D-FRONT
val_unseen
6
5
M3DLayout / Matterport3D
输入:英文使用需求 + 可用地板多边形 + 4–10 件家具及固定宽深尺寸。输出:所有家具的二维位置与旋转角度。家具清单和尺寸不可修改。参考摆放已通过几何检查,但没有被认证为满足全部语言偏好的标准答案。
下载后打开 review.html 可以逐条浏览需求、尺寸、空房轮廓、参考图和修订理由。原文与 46 条逐条审核记录见 individual_reviews.jsonl,其中 39 条保留、7 条排除。此次规模适合跑通… See the full description on the dataset page: https://huggingface.co/datasets/yfan1997/room-layout-planning-curated-v1.docbank-layout
Support the Project ☕
If you find this dataset helpful, please support me with a mocha:
Dataset Summary
DocBank is a large-scale dataset tailored for Document AI tasks, focusing on integrating textual and layout information. It comprises 500,000 document pages, divided into 400,000 for training, 50,000 for validation, and 50,000 for testing. The dataset is generated using a weak supervision approach, enabling efficient annotation of document structures… See the full description on the dataset page: https://huggingface.co/datasets/astrologos/docbank-layout.layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation.
Project page: https://orangesodahub.github.io/SceneCraft
Code: https://github.com/OrangeSodahub/SceneCraft
SORIE_layoutlmv2
Dataset Card for "SORIE_layoutlmv2"
More Information needed
layout_distribution_shiftlayout_reconstruction-preset-gemini
layout_reconstruction-preset-gemini
Real-robot teleoperation episodes of the layout_reconstruction task on a single-arm Franka Research 3 cell (80 episodes, 37,559 frames at 10 fps,
released as Myungkyu/layout_reconstruction) with dense high-level labels produced by the TACOR offline annotator:
Gemini 3.7 Flash reads each whole episode as one video clip (one sample every 10 frames = 1.0 s) and labels every sampled frame given only the
subtask preset of the task - the label list… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/layout_reconstruction-preset-gemini.regulated_layout_dataset_v10_20260830FloorplanQA-Layouts
FloorplanQA (Layouts Only)
This repository contains 2,000 JSON layouts used in the paper:
"FloorplanQA: A Benchmark for Spatial Reasoning in LLMs Using Structured Representations"
arXiv: https://arxiv.org/abs/2507.07644
Project Page: https://olddelorean.github.io/FloorplanQA/
Contents
600 synthetic kitchens
600 synthetic living rooms
600 synthetic bedrooms
200 layouts derived from HSSD-200
The release includes only the layouts.
Structure
layouts/… See the full description on the dataset page: https://huggingface.co/datasets/OldDelorean/FloorplanQA-Layouts.layoutbench
LayoutBench
Release of LayoutBench dataset from Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation (CVPR 2024 Workshop)
See also LayoutBench-COCO for zero-shot evaluation on OOD layouts with real objects.
[Project Page]
[Paper]
Authors:
Jaemin Cho,
Linjie Li,
Zhengyuan Yang,
Zhe Gan,
Lijuan Wang,
Mohit Bansal
Summary
LayoutBench is a diagnostic benchmark that examines layout-guided image generation models on arbitrary, unseen layouts.… See the full description on the dataset page: https://huggingface.co/datasets/j-min/layoutbench.3d_layout_reasoningDataset for MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
Github: https://github.com/PzySeere/MetaSpatial
layout_diffusion_scannetpp_voxel0.2This dataset is used in the paper SceneCraft: Layout-Guided 3D Scene Generation.
File information
The repository contains the following file information:
DocVQA_LayoutLM_features
Dataset Card for "DocVQA_LayoutLM_features"
More Information needed
layouts_procthorpii-layout-synthLayoutOrderingHardlayout_reconstruction
layout_reconstruction
Real-robot teleoperation demonstrations of the layout_reconstruction task on a single-arm Franka Research 3 cell,
released in four LeRobot layouts. Every layout is a conversion of the same 80 raw episodes
(37,559 frames at 10 Hz); the layouts differ only in the LeRobot codebase version and in the
action representation.
directory
LeRobot version
action (action)
consumer
lerobot_v21_abs_joint/
v2.1
8-D absolute joint targets + gripper
RLDX-1 loader… See the full description on the dataset page: https://huggingface.co/datasets/Myungkyu/layout_reconstruction.DocVQA_layoutLM
Dataset Card for "DocVQA_layoutLM"
More Information needed
PCB-Layout
PCB-Layout: Aligned PCB–Layout Pairs
Project Page ·
Online Demo ·
Code
PCB-Layout contains 668 pairs of PCB photographs and circuit layout images, with each pair manually aligned to establish spatial correspondence.
Preview
Rows: PCB · Layout · Animated overlay (PCB ↔ blend ↔ Layout)
Acquisition & Alignment… See the full description on the dataset page: https://huggingface.co/datasets/Warren-wzw/PCB-Layout.persian-ocr-community-dataset-layout
Persian OCR Community Layout Annotations
Resumable layout annotations for the page images in
Reza2kn/persian-ocr-community-dataset.
Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the
page identifier, page dimensions, handwriting flag, and structured layout boxes produced by
datalab-to/surya_layout2 at confidence threshold
0.4.
The boxes field contains label, confidence, raster-order position, and pixel coordinates
x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.LayoutPrompterA collection of datasets used in LayoutPrompter (NeurIPS2023).
Specifically, publaynet and rico are downloaded from LayoutFormer++, posterlayout is downloaded from DS-GAN, and webui is downloaded from Parse-Then-Place.
We sincerely thank them for the great work they do.
