Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SQuADDS /SQuADDS_Layouts SQuADDS Layouts - versioned GDS artifacts for superconducting quantum hardware SQuADDS Layouts is the geometry-artifact companion to SQuADDS_DB, the Superconducting Qubit And Device Design and Simulation Database. It provides checksum-verified GDS files, stable geometry identities, and machine-readable geometry metadata so a simulation result can be traced to the exact layout that produced it. Homepage: https://lfl-lab.github.io/SQuADDS/ Repository:… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layouts.tabular10K<n<100K1 likes5.9k downloads1mo agoHugging Face02HuiZhang0812 /LayoutSAM LayoutSAM Dataset Overview The LayoutSAM dataset is a large-scale layout dataset derived from the SAM dataset, containing 2.7 million image-text pairs and 10.7 million entities. Each entity is annotated with a spatial position (i.e., bounding box) and a textual description. Traditional layout datasets often exhibit a closed-set and coarse-grained nature, which may limit the model's ability to generate complex attributes such as color, shape, and texture.… See the full description on the dataset page: https://huggingface.co/datasets/HuiZhang0812/LayoutSAM.text1M<n<10M11 likes1.2k downloads2y agoHugging Face03nielsr /funsd-layoutlmv3imagen<1K42 likes794 downloads1y agoHugging Face04nakamura196 /ndl-layout-dataset NDL-DocL Kotenseki Layout Dataset (YOLO format) A YOLO-formatted conversion of the kotenseki (pre-modern Japanese materials, 古典籍資料) subset of the NDL-DocL dataset published by the National Diet Library of Japan (NDL). Source dataset: https://github.com/ndl-lab/layout-dataset Source images: NDL Digital Collections https://dl.ndl.go.jp/ 国立国会図書館が公開する NDL-DocL データセットのうち、古典籍資料を YOLO 形式(Ultralytics 互換)に変換したものです。 This is a modified/derived version. The bounding boxes were converted… See the full description on the dataset page: https://huggingface.co/datasets/nakamura196/ndl-layout-dataset.imageobject-detection1K<n<10K1 likes741 downloads23d agoHugging Face05EditFigure /regulated_layout_dataset_v9_20260802text100K<n<1M0 likes733 downloads2mo agoHugging Face06yfan1997 /room-layout-planning-curated-v1 Room Layout Planning — curated pilot v1 39 个逐条检查并编写需求的房间布局任务,供实验流程验证与人工抽查。所有最终设计需求均为 AI 编写;没有人工标注或人工复核声明。 Split 条数 独立房屋 几何来源 train 26 26 InstructScene / 3D-FRONT val_seen 7 7 InstructScene / 3D-FRONT val_unseen 6 5 M3DLayout / Matterport3D 输入:英文使用需求 + 可用地板多边形 + 4–10 件家具及固定宽深尺寸。输出:所有家具的二维位置与旋转角度。家具清单和尺寸不可修改。参考摆放已通过几何检查,但没有被认证为满足全部语言偏好的标准答案。 下载后打开 review.html 可以逐条浏览需求、尺寸、空房轮廓、参考图和修订理由。原文与 46 条逐条审核记录见 individual_reviews.jsonl,其中 39 条保留、7 条排除。此次规模适合跑通… See the full description on the dataset page: https://huggingface.co/datasets/yfan1997/room-layout-planning-curated-v1.imagen<1K0 likes518 downloads18d agoHugging Face07EditFigure /regulated_layout_dataset_v10_20260830text100K<n<1M0 likes461 downloads1mo agoHugging Face08astrologos /docbank-layout Support the Project ☕ If you find this dataset helpful, please support me with a mocha: Dataset Summary DocBank is a large-scale dataset tailored for Document AI tasks, focusing on integrating textual and layout information. It comprises 500,000 document pages, divided into 400,000 for training, 50,000 for validation, and 50,000 for testing. The dataset is generated using a weak supervision approach, enabling efficient annotation of document structures… See the full description on the dataset page: https://huggingface.co/datasets/astrologos/docbank-layout.textgraph-ml100K<n<1M1 likes407 downloads2y agoHugging Face09gzzyyxy /layout_diffusion_scannetpp_voxel0.2This dataset is used in the paper SceneCraft: Layout-Guided 3D Scene Generation. File information The repository contains the following file information: imagetext-to-3d10K<n<100K2 likes365 downloads1y agoHugging Face10Extend-AI /RealDoc-Bench-Layout RealDocBench-Layout A 1,500-page document-layout benchmark for evaluating layout-detection models on real-world documents. COCO-style annotations across 9 block classes. Contents images/ — 1,500 page images (PNG / JPG / occasional WebP-as-PNG; see Caveats). annotations/<pageId>.json — per-page COCO files, each with a single image record, an annotations list, a categories list, and a page_info block. manifest.csv — pageId → domain + source URLs. The canonical row… See the full description on the dataset page: https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout.imageobject-detection1K<n<10K6 likes361 downloads5mo agoHugging Face11OldDelorean /FloorplanQA-Layouts FloorplanQA (Layouts Only) This repository contains 2,000 JSON layouts used in the paper: "FloorplanQA: A Benchmark for Spatial Reasoning in LLMs Using Structured Representations" arXiv: https://arxiv.org/abs/2507.07644 Project Page: https://olddelorean.github.io/FloorplanQA/ Contents 600 synthetic kitchens 600 synthetic living rooms 600 synthetic bedrooms 200 layouts derived from HSSD-200 The release includes only the layouts. Structure layouts/… See the full description on the dataset page: https://huggingface.co/datasets/OldDelorean/FloorplanQA-Layouts.text1K<n<10K5 likes355 downloads5mo agoHugging Face12allenai /layout_distribution_shifttext10K<n<100K0 likes286 downloads3y agoHugging Face13gzzyyxy /layout_diffusion_hypersimThis repository contains the data for SceneCraft: Layout-Guided 3D Scene Generation. Project page: https://orangesodahub.github.io/SceneCraft Code: https://github.com/OrangeSodahub/SceneCraft imagetext-to-3d10K<n<100K1 likes280 downloads1y agoHugging Face14Reza2kn /persian-ocr-community-dataset-layout Persian OCR Community Layout Annotations Resumable layout annotations for the page images in Reza2kn/persian-ocr-community-dataset. Each row points to an exact source dataset revision, Parquet shard, blob, and row. It includes the page identifier, page dimensions, handwriting flag, and structured layout boxes produced by datalab-to/surya_layout2 at confidence threshold 0.4. The boxes field contains label, confidence, raster-order position, and pixel coordinates x0, y0, x1, y1.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-community-dataset-layout.tabularobject-detection10K<n<100K0 likes261 downloads3mo agoHugging Face15Ssunbell /SORIE_layoutlmv2 Dataset Card for "SORIE_layoutlmv2" More Information needed text100K<n<1M0 likes251 downloads4y agoHugging Face16zhenyupan /3d_layout_reasoningDataset for MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse Github: https://github.com/PzySeere/MetaSpatial imageimage-text-to-textn<1K2 likes237 downloads2y agoHugging Face17clementchamalet /sky130a-layout-verification-pex Sky130A Cell Layout Verification Dataset and PEX Benchmark 3,005 verified Sky130A cells for cell-level DRC/LVS studies and post-layout PEX modeling. The bundles provide layouts, independent SPICE references, verification and extraction evidence, and source provenance. The benchmark provides parasitic resistance and capacitance labels, a frozen group-disjoint split, baseline predictions, and scores. A provenance census records taxonomy and release decisions for 11,633 catalogued… See the full description on the dataset page: https://huggingface.co/datasets/clementchamalet/sky130a-layout-verification-pex.text10K<n<100K0 likes214 downloads1d agoHugging Face18Sharka /DocVQA_layoutLM Dataset Card for "DocVQA_layoutLM" More Information needed tabular10K<n<100K0 likes178 downloads3y agoHugging Face19R2aillc /LayoutOrderingHardimage1K<n<10K1 likes175 downloads2y agoHugging Face20dauvannam321 /fineweb-vi-pdf-layout Fineweb VI PDF Layout (v2) Vietnamese web pages (from HuggingFaceFW/fineweb-2) rendered to PDF, with per-page layout annotations for VLM training. Each row = 1 document: Column Type Description doc_name string Document ID (e.g. doc_000) url string Source URL (metadata) pdf binary Rendered document.pdf layouts_merged string All per-page layout JSONs merged: {"num_pages": N, "pages": [...]} pages image list Rendered page images (page_*.png) source_html string… See the full description on the dataset page: https://huggingface.co/datasets/dauvannam321/fineweb-vi-pdf-layout.imageimage-to-text1K<n<10K0 likes160 downloads19d agoHugging Face21HuiZhang0812 /LayoutSAM-eval LayoutSAM-eval Benchmark Overview LayoutSAM-Eval is a comprehensive benchmark for evaluating the quality of Layout-to-Image (L2I) generation models. This benchmark assesses L2I generation quality from two perspectives: region-wise quality (spatial and attribute accuracy) and global-wise quality (visual quality and prompt following). It employs the VLM’s visual question answering to evaluate spatial and attribute adherence, and utilizes various metrics including IR score… See the full description on the dataset page: https://huggingface.co/datasets/HuiZhang0812/LayoutSAM-eval.image1K<n<10K2 likes159 downloads2y agoHugging Face22openfoodfacts /nutrient-detection-layout Nutrient extraction dataset This dataset contains annotated images of nutrition tables. The goal of this dataset was to train a model to extract nutrient values from nutrition tables, as part of the Nutrisight project. It contains ~3k samples in total (2.8k for training and 199 for testing). For more information about the project, please refer to the nutrisight directory in the openfoodfacts-ai GitHub repository. The images were collected from the Open Food Facts database, and… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/nutrient-detection-layout.imagetoken-classification1K<n<10K4 likes139 downloads2y agoHugging Face23hydroshiba /hcmus-doc-layout HCMUS Document-Layout Detection Fine-grained document-layout object detection on thesis pages: real scanned HCMUS (Vietnamese) theses plus programmatically generated synthetic pages. Layout (mirrored under the repo root) Path Contents images/train/shard00 … shard03 26,988 JPEG (20,292 unique pages: 6,696 human-reviewed real + 13,596 synthetic; each real page replayed twice as rep2_* aliases → effective ~1:1 real:synthetic per epoch) images/val/ 916… See the full description on the dataset page: https://huggingface.co/datasets/hydroshiba/hcmus-doc-layout.imageobject-detection10K<n<100K0 likes128 downloads1mo agoHugging Face24SpringRollMonster /CIVQA-TesseractOCR-LayoutLM CIVQA TesseractOCR LayoutLM Dataset The Czech Invoice Visual Question Answering dataset was created with Tesseract OCR and encoded for the LayoutLM. The pre-encoded dataset can be found on this link: https://huggingface.co/datasets/fimu-docproc-research/CIVQA-TesseractOCR All invoices used in this dataset were obtained from public sources. Over these invoices, we were focusing on 15 different entities, which are crucial for processing the invoices. Invoice number Variable… See the full description on the dataset page: https://huggingface.co/datasets/SpringRollMonster/CIVQA-TesseractOCR-LayoutLM.tabular100K<n<1M0 likes125 downloads2mo agoHugging Face25SQuADDS /SQuADDS_Layout_Embeddings SQuADDS Layout Embeddings Versioned layout representations for the 24,106 GDS artifacts in SQuADDS/SQuADDS_Layouts. Static embedding model v0 static-embedding-v0 implements the original SQuADDS proof-of-concept model: v0 = parameter_sum + geometric_moments + flattened_shape_bitmap Each unit-normalized vector has 9,227 dimensions: Block Dimensions Contents Parameter sum 1 Permutation- and parameter-count-invariant sum of numerical design options… See the full description on the dataset page: https://huggingface.co/datasets/SQuADDS/SQuADDS_Layout_Embeddings.tabular10K<n<100K1 likes120 downloads1mo agoHugging Face26raphael0202 /ingredient-detection-layout-dataset Dataset Card for "ingredient-detection-layout-dataset" More Information needed image1K<n<10K0 likes118 downloads3y agoHugging Face27jonny122 /khmer-newspaper-layout-dataset Khmer Newspaper Layout Dataset Dataset Description This dataset contains Khmer newspaper layouts with annotated regions for document layout analysis and OCR tasks. Dataset Summary Total Examples: 9,344 newspaper layouts Language: Khmer (Cambodian) Image Format: PNG Annotations: LabelMe JSON format with bounding boxes and segmentation masks Features config_id: Unique identifier for each sample image: Newspaper layout image (PNG)… See the full description on the dataset page: https://huggingface.co/datasets/jonny122/khmer-newspaper-layout-dataset.imageobject-detection1K<n<10K1 likes118 downloads8mo agoHugging Face28EditFigure /regulated_layout_dataset_v8_20260709text100K<n<1M0 likes114 downloads3mo agoHugging Face29Axi404 /GenManip-Banana-Layoutstextn<1K0 likes87 downloads1y agoHugging Face30NealCaren /co-newspaper-layout-yoloimage1K<n<10K0 likes83 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.