hydroshiba/hcmus-doc-layout
HCMUS Document-Layout Detection Fine-grained document-layout object detection on thesis pages: real scanned HCMUS (Vietnamese) theses plus programmatically generated synthetic pages. Layout (mirrored under the repo root) Path Contents images/train/shard00 … shard03 26,988 JPEG (20,292 unique pages: 6,696 human-reviewed real + 13,596 synthetic; each real page replayed twice as rep2_* aliases → effective ~1:1 real:synthetic per epoch) images/val/ 916… See the full description on the dataset page: https://huggingface.co/datasets/hydroshiba/hcmus-doc-layout.
HCMUS Document-Layout Detection
Fine-grained document-layout object detection on thesis pages: real scanned HCMUS (Vietnamese) theses plus programmatically generated synthetic pages.
Layout (mirrored under the repo root)
Directories are sharded because the Hub allows ≤ 10,000 files per directory.
Classes (15)
cap, catalogue, code, equ, fig, foot, header, list, para, para_equ, reference, sec_1, sec_2, sec_3, tab
Formats — both annotation kinds provided
- COCO
train.json/val.json— images / annotations / categories. Categoryid= line index in the class list above (0-based); boxes in absolute pixels (x y w h);iscrowd: 0.file_nameis repo-relative, e.g.images/train/shard00/HCMUS_2020_..._p001.jpg. - YOLO txt under
annotations/...— one line per box, same filename as the image,class_id cx cy w hnormalized.
Images are 896×896.
Real vs synthetic
- Real: scanned human-reviewed HCMUS theses, split by author so a thesis never straddles train/val (shared fonts/margins/header style would otherwise leak near-duplicates across the split).
- Synthetic: LaTeX-generated pages targeting two corpus weaknesses — severe class imbalance (e.g.
codeis ~0.2% of boxes) and single-genre narrowness. Boxes are exact by construction.
Notes
- The validation set is deliberately real-only: a metric measured partly on generated pages would report how well the model learned the generator rather than how well it reads documents.
- Training (and the reference detector) used this train set as-is, replay included —
rep2_*entries duplicate real pages and are not novel data. test/and benchmark splits are intentionally not published here.
