Team Ai
Datasetpublic

hydroshiba/hcmus-doc-layout

HCMUS Document-Layout Detection Fine-grained document-layout object detection on thesis pages: real scanned HCMUS (Vietnamese) theses plus programmatically generated synthetic pages. Layout (mirrored under the repo root) Path Contents images/train/shard00 … shard03 26,988 JPEG (20,292 unique pages: 6,696 human-reviewed real + 13,596 synthetic; each real page replayed twice as rep2_* aliases → effective ~1:1 real:synthetic per epoch) images/val/ 916… See the full description on the dataset page: https://huggingface.co/datasets/hydroshiba/hcmus-doc-layout.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes128downloads
Dataset Card

HCMUS Document-Layout Detection

Fine-grained document-layout object detection on thesis pages: real scanned HCMUS (Vietnamese) theses plus programmatically generated synthetic pages.

Layout (mirrored under the repo root)

PathContents
images/train/shard00 … shard0326,988 JPEG (20,292 unique pages: 6,696 human-reviewed real + 13,596 synthetic; each real page replayed twice as rep2_* aliases → effective ~1:1 real:synthetic per epoch)
images/val/916 JPEG, real only, held out at the author/document level
annotations/train/shard00 … shard03YOLO txt, same shard/filename as the images
annotations/val/YOLO txt
train.json, val.jsonCOCO annotations

Directories are sharded because the Hub allows ≤ 10,000 files per directory.

Classes (15)

cap, catalogue, code, equ, fig, foot, header, list, para, para_equ, reference, sec_1, sec_2, sec_3, tab

Formats — both annotation kinds provided

  1. 1.COCO train.json / val.json — images / annotations / categories. Category id = line index in the class list above (0-based); boxes in absolute pixels (x y w h); iscrowd: 0. file_name is repo-relative, e.g. images/train/shard00/HCMUS_2020_..._p001.jpg.
  2. 2.YOLO txt under annotations/... — one line per box, same filename as the image, class_id cx cy w h normalized.

Images are 896×896.

Real vs synthetic

  • —Real: scanned human-reviewed HCMUS theses, split by author so a thesis never straddles train/val (shared fonts/margins/header style would otherwise leak near-duplicates across the split).
  • —Synthetic: LaTeX-generated pages targeting two corpus weaknesses — severe class imbalance (e.g. code is ~0.2% of boxes) and single-genre narrowness. Boxes are exact by construction.

Notes

  • —The validation set is deliberately real-only: a metric measured partly on generated pages would report how well the model learned the generator rather than how well it reads documents.
  • —Training (and the reference detector) used this train set as-is, replay included — rep2_* entries duplicate real pages and are not novel data.
  • —test/ and benchmark splits are intentionally not published here.
hydroshiba/hcmus-doc-layout · Team Ai