Team Ai
Datasetpublic

kbang2021/doclaynet-6class

DocLayNet 6-Class Filtered Dataset Dataset Description This is a filtered version of the DocLayNet dataset containing only 6 most relevant layout element classes for document layout analysis tasks. Original Dataset DocLayNet is a human-annotated document layout segmentation dataset containing 80,863 pages from diverse sources with 11 distinct layout categories. Citation: @article{doclaynet2022, title = {DocLayNet: A Large Human-Annotated Dataset… See the full description on the dataset page: https://huggingface.co/datasets/kbang2021/doclaynet-6class.

sourceHugging Facecdla-permissive-2.0updated 7mo agoView on Hugging Face
0likes22downloads
Dataset Card

DocLayNet 6-Class Filtered Dataset

Dataset Description

This is a filtered version of the DocLayNet dataset containing only 6 most relevant layout element classes for document layout analysis tasks.

Original Dataset

DocLayNet is a human-annotated document layout segmentation dataset containing 80,863 pages from diverse sources with 11 distinct layout categories.

Citation:

@article{doclaynet2022,
  title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis},
  author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S and Staar, Peter W J},
  year = {2022},
  doi = {10.1145/3534678.3539043},
}

Filtering Methodology

Classes Retained (6):

  1. 1.Text - Body text paragraphs
  2. 2.List-item - List elements (bulleted, numbered)
  3. 3.Section-header - Section and subsection titles
  4. 4.Picture - Images, figures, diagrams
  5. 5.Table - Tabular data structures
  6. 6.Caption - Image and table captions

Classes Removed (5):

  • —Footnote
  • —Formula
  • —Page-footer
  • —Page-header
  • —Title

Rationale: Focus on the most common and semantically important layout elements for general document understanding tasks. The 6 retained classes represent 85.1% of all annotations in the original dataset.

Dataset Statistics

Split Distribution

SplitImagesAnnotationsClasses
Train68,673800,6146
Validation6,44685,0576
Test4,95256,4836
Total80,071942,1546

Class Distribution (Training Set)

Based on 800,614 annotations:

Class IDClass NameCountPercentage
0Caption19,2182.4%
1List-item161,81820.2%
2Picture39,6675.0%
3Section-header118,59014.8%
4Table30,0703.8%
5Text431,25153.9%

Retention from Original Dataset

  • —Images retained: 99.0%
  • —Annotations retained: 85.1%

Dataset Structure

Format

Annotations are provided in COCO JSON format:

DocLayNet_6class/
├── coco/
│   ├── train.json      # Training annotations
│   ├── val.json        # Validation annotations
│   └── test.json       # Test annotations
└── README.md           # This file

Images are NOT included - use the original DocLayNet image files from:

  • —HuggingFace: docling-project/DocLayNet
  • —Official source: https://github.com/DS4SD/DocLayNet

Loading the Dataset

Using HuggingFace Datasets
python
from datasets import load_dataset

# Load the filtered annotations
dataset = load_dataset("kbang2021/doclaynet-6class")

# Access splits
train_data = dataset["train"]
val_data = dataset["validation"]
test_data = dataset["test"]
Manual Loading
python
import json
from pathlib import Path

# Load COCO annotations
with open("coco/train.json") as f:
    train_coco = json.load(f)

# Categories
categories = train_coco["categories"]  # 6 classes with IDs 0-5

# Images
images = train_coco["images"]  # Image metadata

# Annotations
annotations = train_coco["annotations"]  # Bounding boxes

Annotation Format

Each annotation follows the COCO format:

json
{
  "id": 12345,
  "image_id": 123,
  "category_id": 5,  // 0-5 (remapped from original 11 classes)
  "bbox": [x_min, y_min, width, height],  // In pixels
  "area": 12345.67,
  "iscrowd": 0
}

Category Mapping

Original DocLayNet → 6-Class Filtered:

Original IDOriginal NameFiltered IDFiltered NameStatus
0Caption0Caption✅ Kept
1Footnote--❌ Removed
2Formula--❌ Removed
3List-item1List-item✅ Kept
4Page-footer--❌ Removed
5Page-header--❌ Removed
6Picture2Picture✅ Kept
7Section-header3Section-header✅ Kept
8Table4Table✅ Kept
9Text5Text✅ Kept
10Title--❌ Removed

Use Cases

This filtered dataset is ideal for:

  • —Document layout analysis with focus on content structure
  • —Information extraction from documents (text, tables, figures)
  • —Object detection model training for document AI
  • —Multi-scale document understanding tasks
  • —Transfer learning from general object detection to document analysis

Limitations

  1. 1.Images not included: You must obtain images from the original DocLayNet dataset
  2. 2.Class imbalance: Text class dominates (53.9% of annotations)
  3. 3.Domain specific: Focused on document layout, may not generalize to other domains
  4. 4.Annotation quality: Inherits any annotation errors from original DocLayNet

Ethical Considerations

  • —Dataset maintains the original DocLayNet license (CDLA-Permissive-2.0)
  • —No personal or sensitive information in annotations
  • —Source documents from diverse domains (financial, scientific, patents, manuals)
  • —Should not be used to discriminate based on document type or origin

Citation

If you use this filtered dataset, please cite both:

  1. 1.Original DocLayNet paper:
bibtex
@article{doclaynet2022,
  title = {DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis},
  author = {Pfitzmann, Birgit and Auer, Christoph and Dolfi, Michele and Nassar, Ahmed S and Staar, Peter W J},
  year = {2022},
  doi = {10.1145/3534678.3539043},
}
  1. 1.This filtered version:
bibtex
@misc{doclaynet6class2024,
  title = {DocLayNet 6-Class: Filtered Document Layout Analysis Dataset},
  author = {[Keng Boon, Ang]},
  year = {2026},
  howpublished = {\url{https://huggingface.co/datasets/kbang2021/doclaynet-6class}},
  note = {Filtered subset of DocLayNet containing 6 primary layout element classes}
}

License

This filtered dataset maintains the original license:

CDLA-Permissive-2.0 (Community Data License Agreement – Permissive – Version 2.0)

See: https://cdla.dev/permissive-2-0/

Acknowledgments

  • —Original DocLayNet dataset: IBM Research
  • —Built using the layout-for-tools evaluation framework

Contact

For questions or issues with this filtered dataset, please open an issue on the repository.

For questions about the original DocLayNet dataset, see: https://github.com/DS4SD/DocLayNet