freococo/myanmar_complex_document_layouts
๐ฒ๐ฒ Myanmar Complex Document Layouts A large-scale, high-quality synthetic dataset containing 17,632 images of complex document layouts, dashboards, and infographics entirely in the Myanmar (Burmese) language. This dataset is specifically designed to train and benchmark modern Computer Vision and multimodal LLMs on complex Myanmar typography, structured data, and diverse graphical layouts. ๐ Dataset Overview Total Images: 17,632 high-resolution pages.โฆ See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_complex_document_layouts.
๐ฒ๐ฒ Myanmar Complex Document Layouts
A large-scale, high-quality synthetic dataset containing 17,632 images of complex document layouts, dashboards, and infographics entirely in the Myanmar (Burmese) language.
This dataset is specifically designed to train and benchmark modern Computer Vision and multimodal LLMs on complex Myanmar typography, structured data, and diverse graphical layouts.
๐ Dataset Overview
- Total Images: 17,632 high-resolution pages.
- Language: Myanmar (Burmese).
- Format: Parquet (Optimized for streaming and fast loading).
- Content: Dictionary definitions, part-of-speech (POS) tagging, phonetics, related word tables, and dynamic graphical charts.
๐ Key Features
- Rich Typography: Utilizes a diverse pool of Myanmar fonts classified into Body, Bold, UI, and Decorative styles.
- Complex Layouts: Alternating 60/40 and 40/60 split-grid layouts featuring multi-line paragraphs, structured data tables, and metric summary boxes.
- Diverse Data Visualizations: Includes 9 different chart categories (Column, Bar, Line, Area, Pie, Donut, Radar, Scatter, Bubble) rendered in 2D, 3D, and Interactive styles.
- Perfect Bounding Boxes: The metadata provides exact pixel-perfect bounding box coordinates (
[left, top, right, bottom]) for every text element, chart wrapper, and table cell on the page. - Realistic Dimensions: Rendered in standard A4 and US Letter sizes, covering both Portrait and Landscape orientations.
๐ Primary Use Cases
- Document OCR & Text Extraction: Training models to accurately read heavily formatted Myanmar text without spacing issues.
- Document Layout Analysis: Training models like LayoutLM, Donut, or Docling to understand reading order, columns, and spatial relationships.
- Document Visual Question Answering (DocVQA): Benchmarking multimodal LLMs on their ability to extract facts from Myanmar tables and metric boxes.
- Chart Understanding: Teaching AI to interpret visual data structures associated with Myanmar labels.
๐๏ธ Metadata Structure
Each row in the dataset provides the rendered image alongside rich JSON metadata, including:
page_number&page_size_namelayout_style&themecolorsfonts_used(Specific TTF files used for that page)chart_1_data&chart_2_data(The raw numerical data backing the charts)entries: A list of all extracted text strings and their exact bounding boxes.table_bbox,profile_card_bbox,chart_bboxes: Bounding boxes for the major UI containers.
โ๏ธ Creation Process
This dataset was procedurally generated using Python and a headless Chromium browser engine (Playwright). Dynamic CSS Grid styling, SVG generation, and strict line-breaking rules were utilized to ensure the text and graphics realistically represent modern digital documents and infographics.
