Team Ai
Datasetpublic

mapo80/DocCornerDataset

DocCornerDataset A high-quality document corner detection dataset for training models to detect the four corners of documents in images. This dataset is optimized for building robust document scanning and perspective correction applications. Dataset Examples Training Set Validation Set Test Set Green polygons show the annotated document corners Dataset Description This dataset contains images with… See the full description on the dataset page: https://huggingface.co/datasets/mapo80/DocCornerDataset.

sourceHugging Faceotherupdated 9mo agoView on Hugging Face
7likes1.1kdownloads
Dataset Card

DocCornerDataset

A high-quality document corner detection dataset for training models to detect the four corners of documents in images. This dataset is optimized for building robust document scanning and perspective correction applications.

Dataset Examples

Training Set

<img src="collages/train_collage.jpg" alt="Training samples" width="600"/>

Validation Set

<img src="collages/val_collage.jpg" alt="Validation samples" width="600"/>

Test Set

<img src="collages/test_collage.jpg" alt="Test samples" width="600"/>

Green polygons show the annotated document corners

Dataset Description

This dataset contains images with document corner annotations, optimized for training robust document detection models. It uses the best-performing splits from an iterative dataset cleaning process with multiple quality validation steps.

Key Features

  • —High Quality Annotations: Labels refined through iterative cleaning with multiple teacher models
  • —Diverse Document Types: IDs, invoices, receipts, books, cards, and general documents
  • —Negative Samples: Includes images without documents for training robust classifiers
  • —No Overlap: Train, validation, and test splits are completely disjoint

Dataset Statistics

SplitImagesDescription
train32,968Training set (cleaned iter3 + hard negatives)
validation8,645Validation set (cleaned iter3)
test6,652Held-out test set (no overlap with train/val)
Total48,265

Data Sources and Licenses

This dataset is compiled from multiple open-source datasets. Please refer to the original dataset licenses before using this data.

MIDV Dataset (ID Cards)

Mobile Identity Document Video dataset for identity document detection and recognition.

DatasetImagesLicenseSource
MIDV-500~9,400Research useWebsite
MIDV-2019~1,350Research useWebsite

Citation:

bibtex
@article{arlazarov2019midv500,
  title={MIDV-500: A Dataset for Identity Documents Analysis and Recognition on Mobile Devices in Video Stream},
  author={Arlazarov, V.V. and Bulatov, K. and Chernov, T. and Arlazarov, V.L.},
  journal={Computer Optics},
  volume={43},
  number={5},
  pages={818--824},
  year={2019}
}

@inproceedings{arlazarov2019midv2019,
  title={MIDV-2019: Challenges of the modern mobile-based document OCR},
  author={Arlazarov, V.V. and Bulatov, K. and Chernov, T. and Arlazarov, V.L.},
  booktitle={ICDAR},
  year={2019}
}

SmartDoc Dataset (Documents)

SmartDoc Challenge dataset for document image acquisition and quality assessment.

DatasetImagesLicenseSource
SmartDoc~1,380Research useWebsite

Citation:

bibtex
@inproceedings{burie2015smartdoc,
  title={ICDAR 2015 Competition on Smartphone Document Capture and OCR (SmartDoc)},
  author={Burie, J.C. and Chazalon, J. and Coustaty, M. and others},
  booktitle={ICDAR},
  year={2015}
}

COCO Dataset (Negative Samples)

Common Objects in Context dataset used for negative samples (images without documents).

DatasetImagesLicenseSource
COCO val2017~4,300CC BY 4.0Website
COCO train2017~11,400CC BY 4.0Website

Note: Excluded categories that could be confused with documents: book, laptop, tv, cell phone, keyboard, mouse, remote, clock.

Citation:

bibtex
@inproceedings{lin2014coco,
  title={Microsoft COCO: Common Objects in Context},
  author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and others},
  booktitle={ECCV},
  year={2014}
}

Roboflow Universe (Various Documents)

Various document datasets from Roboflow Universe community.

CategoryDatasetsLicenseSource
Documentsdocumentsegmentationv2, docscanner, docrida, documentoVarious (check individual)Roboflow Universe
Bills/Invoicesbillsegmentation, csinvoiceVarious (check individual)Roboflow Universe
Receiptsreceiptdetection, receiptoccam, receipts_coolstuffVarious (check individual)Roboflow Universe
ID Cardscardcorner, card4class, idcardskew, iddetections, idcard_jjVarious (check individual)Roboflow Universe
Passportssegment_passportVarious (check individual)Roboflow Universe
Booksbookreader, pagesegmentationtecgp, bookcmjt2Various (check individual)Roboflow Universe

Note: Roboflow datasets have various licenses. Please check the individual dataset pages on Roboflow Universe for specific license terms.

Features

FeatureTypeDescription
imageImageThe document image (JPEG)
filenamestringOriginal filename for traceability
is_negativeboolTrue if image contains no document
corner_tl_xfloat32Top-left corner X coordinate (normalized 0-1)
corner_tl_yfloat32Top-left corner Y coordinate (normalized 0-1)
corner_tr_xfloat32Top-right corner X coordinate (normalized 0-1)
corner_tr_yfloat32Top-right corner Y coordinate (normalized 0-1)
corner_br_xfloat32Bottom-right corner X coordinate (normalized 0-1)
corner_br_yfloat32Bottom-right corner Y coordinate (normalized 0-1)
corner_bl_xfloat32Bottom-left corner X coordinate (normalized 0-1)
corner_bl_yfloat32Bottom-left corner Y coordinate (normalized 0-1)

Corner Order

Corners are ordered clockwise starting from top-left:

1 (TL) -------- 2 (TR)
   |              |
   |   Document   |
   |              |
4 (BL) -------- 3 (BR)

Coordinate System

  • —Coordinates are normalized to the range [0, 1]
  • —To convert to pixel coordinates: pixel_x = corner_x * image_width
  • —Origin (0, 0) is at the top-left of the image

Negative Samples

Images with is_negative=True:

  • —Do not contain any document
  • —All corner coordinates are null
  • —Useful for training classifiers to reject non-document images

Usage

Loading the Dataset

python
from datasets import load_dataset

# Load all splits
dataset = load_dataset("mapo80/DocCornerDataset")

# Access specific splits
train_data = dataset["train"]
val_data = dataset["validation"]
test_data = dataset["test"]

print(f"Train: {len(train_data)} samples")
print(f"Val: {len(val_data)} samples")
print(f"Test: {len(test_data)} samples")

Iterating Over Samples

python
for sample in dataset["train"]:
    image = sample["image"]  # PIL Image
    filename = sample["filename"]

    if not sample["is_negative"]:
        # Get corner coordinates (normalized 0-1)
        corners = [
            (sample["corner_tl_x"], sample["corner_tl_y"]),
            (sample["corner_tr_x"], sample["corner_tr_y"]),
            (sample["corner_br_x"], sample["corner_br_y"]),
            (sample["corner_bl_x"], sample["corner_bl_y"]),
        ]

        # Convert to pixel coordinates
        w, h = image.size
        corners_px = [(int(x * w), int(y * h)) for x, y in corners]

Visualizing Annotations

python
from PIL import Image, ImageDraw

def draw_corners(image, corners, color=(0, 255, 0), width=3):
    """Draw document corners on image."""
    draw = ImageDraw.Draw(image)
    w, h = image.size

    # Convert normalized to pixel coords
    points = [(int(c[0] * w), int(c[1] * h)) for c in corners]

    # Draw polygon
    for i in range(4):
        draw.line([points[i], points[(i+1) % 4]], fill=color, width=width)

    # Draw corner circles
    for p in points:
        r = 5
        draw.ellipse([p[0]-r, p[1]-r, p[0]+r, p[1]+r], fill=color)

    return image

# Example usage
sample = dataset["train"][0]
if not sample["is_negative"]:
    corners = [
        (sample["corner_tl_x"], sample["corner_tl_y"]),
        (sample["corner_tr_x"], sample["corner_tr_y"]),
        (sample["corner_br_x"], sample["corner_br_y"]),
        (sample["corner_bl_x"], sample["corner_bl_y"]),
    ]
    annotated = draw_corners(sample["image"].copy(), corners)
    annotated.show()

Training a Model (PyTorch Example)

python
import torch
from torch.utils.data import DataLoader
from datasets import load_dataset

dataset = load_dataset("mapo80/DocCornerDataset")

def collate_fn(batch):
    images = torch.stack([transform(s["image"]) for s in batch])

    # Stack corner coordinates (8 values per sample)
    corners = []
    for s in batch:
        if s["is_negative"]:
            corners.append(torch.zeros(8))
        else:
            corners.append(torch.tensor([
                s["corner_tl_x"], s["corner_tl_y"],
                s["corner_tr_x"], s["corner_tr_y"],
                s["corner_br_x"], s["corner_br_y"],
                s["corner_bl_x"], s["corner_bl_y"],
            ]))

    return images, torch.stack(corners)

train_loader = DataLoader(
    dataset["train"],
    batch_size=32,
    shuffle=True,
    collate_fn=collate_fn
)

Model Performance

Models trained on this dataset achieve the following performance:

ModelInput SizemIoU (val)mIoU (test)
MobileNetV2 (alpha=0.35)224x2240.98940.9826
MobileNetV2 (alpha=0.35)256x2560.99020.9819

mIoU = Mean Intersection over Union between predicted and ground truth quadrilaterals

Citation

If you use this dataset in your research, please cite this dataset and the original source datasets:

bibtex
@dataset{doccornerdataset2025,
  author = {mapo80},
  title = {DocCornerDataset: Document Corner Detection Dataset},
  year = {2025},
  publisher = {Hugging Face},
  url = {https://huggingface.co/datasets/mapo80/DocCornerDataset}
}

Please also cite the original datasets used:

  • —MIDV-500/MIDV-2019 (Arlazarov et al., 2019)
  • —SmartDoc (Burie et al., 2015)
  • —COCO (Lin et al., 2014)

License

⚠️ This dataset is compiled from multiple sources with different licenses.

SourceLicense
MIDV-500/MIDV-2019Research use only
SmartDocResearch use only
COCOCC BY 4.0
Roboflow datasetsVarious (check individual datasets)

Before using this dataset, please review the licenses of the original datasets:

Acknowledgments

This dataset was created by combining and processing multiple open-source datasets. We thank the authors of MIDV, SmartDoc, COCO, and the Roboflow community for making their data available.

Related Projects

  • —DocCornerNet - Document corner detection model trained on this dataset