Darayut/khmer-text-recognition
<div align="center"> <img src="assets/netra-logo-transparent-new.png" width="60%" alt="Netra Lab" /> </div>
<hr>
<p align="center"> <a href="https://github.com/netra-ai-lab/Netra-OCR"><b>GitHub</b></a> | <a href="https://pypi.org/project/netra-ocr/"><b>PyPI</b></a> | <a href="https://github.com/netra-ai-lab/Netra-OCR#netra-ocr-app-docker"><b>Desktop app (Docker)</b></a> | <a href="https://huggingface.co/collections/Darayut/khmer-text-synthetic"><b>Datasets</b></a> | <a href="https://huggingface.co/spaces/Darayut/Khmer-Text-Recognition"><b>Inference Space</b></a> </p>
<h2> <p align="center"> A Squeeze-and-Excitation Transformer Network for Khmer Optical Character Recognition </p> </h2>
<p align="center"> <img src="assets/benchmark.png" style="width: 1000px" align=center> </p>
<p align="center"> Character Error Rate (CER %) on KHOB, Legal Documents and Printed Words. <i>Lower is better.</i> </p>
Introduction
Netra-OCR is a 17M-parameter text-line recognizer for Khmer and English, trained from scratch on 1.3M synthetic text-line images. A Squeeze-and-Excitation network and a Transformer encoder read the image; a Transformer decoder writes the text. Khmer is decoded as grapheme clusters rather than single code points, which makes Khmer output sequences about 43% shorter than character-level decoding.
The model reads one text line at a time. To read whole pages (scans, phone photos, PDFs), use the `netra-ocr` package, which adds text-line detection, page layout analysis and export to Word, HTML, Markdown and JSON, or run the ready-made Netra OCR app with Docker (below).
Files in this repository
The cluster model's vocabulary (char2idx_cluster.json) ships inside the netra-ocr package, which downloads the checkpoint from this repository automatically.
Usage
Recognize text lines (Python)
pip install netra-ocrfrom netra_ocr.recognition import recognize, recognize_batch
# One cropped text line: a file path or a PIL.Image
text = recognize("line_crop.png")
# Many crops at once; the model stays loaded between calls
texts = recognize_batch(line_crops, batch_size=8)
# Blockwise decoder: same text, usually faster
text = recognize("line_crop.png", decoder="blockwise")
# Beam search: slower, slightly more accurate (autoregressive decoder only)
text = recognize("line_crop.png", beam_width=3)The checkpoint is downloaded from this repository on first use and cached.
Read whole documents (Python)
pip install "netra-ocr[yolo,layout,pdf,docx]"from netra_ocr.ocr_engine import KhmerOCRPipeline
from netra_ocr.exporters import save_document
pipeline = KhmerOCRPipeline(detector="yolo", layout=True)
doc = pipeline.process_document("letter.pdf") # images, multi-page TIFF or PDF
save_document(doc, "letter.docx") # .docx .html .md .json .txtThis detects the text lines, analyses the page layout (headings, paragraphs, tables, figures, headers and footers, reading order) with DocLayout-YOLO, and recognizes each line with this model.
Netra OCR app (no Python needed)
A web app and REST API with every model included. It runs on your own computer, on the CPU, with no internet connection needed after the download:
docker run -p 8000:8000 -v netra-data:/data ghcr.io/netra-ai-lab/netra-ocr
# then open http://localhost:8000Upload a scan, a phone photo or a PDF, correct the recognized document block by block next to the page image, and export it to Word. Works on Windows, Linux and Mac (Intel and Apple Silicon). See the GitHub README for details.
Datasets
The model was trained entirely on synthetic data and evaluated on real-world and synthetic data.
Training data (synthetic, 1.3M text lines)
Evaluation data
Methodology & Architecture
Preprocessing: chunking and merging
To handle variable-length text lines without aggressive resizing, each image is resized to a height of 48 px (keeping its aspect ratio) and split into overlapping 48×100 px chunks with a 16 px overlap.
Model architecture
<p><em>The input image is resized and split into 48×100 px chunks with 16 px overlaps. Each chunk passes through the Squeeze-and-Excitation network in parallel, producing 512 feature maps of 2×32 px, which become patch embeddings with positional embeddings. The Transformer encoder turns them into vision tokens; the tokens of all chunks are merged, smoothed by a bidirectional LSTM, and passed to the Transformer decoder, which outputs the text one character cluster at a time.</em></p>
- Squeeze-and-Excitation network. Five VGG-style convolution blocks (64 → 512 channels) with height-only pooling in the later stages, so the horizontal axis, where character order lives, is preserved. Blocks 3, 4 and 5 add a 1D Squeeze-and-Excitation module that averages over height only, giving each horizontal column its own channel weights, so background and noise can be suppressed independently at each position along the line.
- Patch module. Collapses the feature map's height and projects each column to a 384-dimensional embedding, plus a learnable positional embedding within the chunk.
- Transformer encoder. Self-attention among the patches of a chunk resolves local ambiguities, such as visually similar sub-consonant stacks.
- Merging module. Concatenates the vision tokens of all chunks of a line into one sequence and adds a second, global positional embedding across the whole line.
- BiLSTM context smoother. A bidirectional LSTM runs over the merged sequence so information flows across chunk boundaries in both directions, smoothing the seams where a character is split between two chunks.
- Transformer decoder. Generates the output sequence autoregressively, with cross-attention over the smoothed encoder output.
Tokenization
Latin text is decoded character by character. Khmer text is decoded as Khmer Character Clusters (a consonant with its subscripts and vowels, e.g. ក្រ), with any cluster outside the vocabulary falling back to single characters. On a 35,000-line Khmer corpus this shortens sequences by 42.65% with lossless round-tripping, which also means fewer decoding steps per line.
Blockwise-parallel decoding
khmerocr_cluster_blockwise.pth adds a small proposal head (Stern et al. 2018) on top of the frozen autoregressive model. Each step it proposes several tokens ahead and keeps only those that the base model itself would have produced, so the output is provably identical to greedy `ar` decoding, in fewer steps.
Training
Adam (learning rate 1e-4) with a staged cyclic learning-rate schedule, cross-entropy loss, and 50,000 randomly sampled, augmented lines per epoch. The cluster-vocabulary model was warm-started from the earlier character-level model (vision layers and shared token embeddings transferred), then trained on the full 1.3M-line set.
Evaluation
Character Error Rate (%), lower is better. † Qwen2.5-VL and DeepSeek-OCR were fine-tuned with unsloth on the same training set as Netra-OCR before evaluation.
Decoders
The table above is the benchmark evaluation. This one compares the two decoders of the released checkpoints with each other, under identical settings: the 325 KHOB lines, greedy decoding, an NVIDIA RTX 3060.
The blockwise decoder accepts on average 2.75 tokens per step (out of 4) and is 1.55× faster end to end, with identical output. Beam search (beam_width=3) lowers KHOB CER to about 0.9% at a higher cost per line.
Examples
<p align="center"> <img src="assets/eval1.png" style="width: 1000px" align=center> <img src="assets/eval2.png" style="width: 1000px" align=center> </p>
Limitations
- The model recognizes single text lines. Whole pages need a line detector first; the
netra-ocrpackage includes one. - It was trained on printed text; handwriting was not part of the training data.
- Only Khmer and English are supported.
License
The code and model weights are released under the MIT license. Note that part of the training data, KhmerSynthetic1M, is licensed by its authors for research and academic use only; check its terms before using this model commercially.
References
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. ICLR 2021. arXiv:2010.11929
- TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, Furu Wei. AAAI 2023. arXiv:2109.10282
- Toward a Low-Resource Non-Latin-Complete Baseline: An Exploration of Khmer Optical Character Recognition R. Buoy, M. Iwamura, S. Srun and K. Kise. IEEE Access, vol. 11, pp. 128044–128060, 2023. DOI: 10.1109/ACCESS.2023.3332361
- Squeeze-and-Excitation Networks Jie Hu, Li Shen, and Gang Sun. CVPR 2018. arXiv:1709.01507
- Bidirectional Recurrent Neural Networks Mike Schuster and Kuldip K. Paliwal. IEEE Transactions on Signal Processing, 1997. DOI: 10.1109/78.650093
- Blockwise Parallel Decoding for Deep Autoregressive Models Mitchell Stern, Noam Shazeer, Jakob Uszkoreit. NeurIPS 2018. arXiv:1811.03115
- DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception Zhiyuan Zhao, Hengrui Kang, Bin Wang, Conghui He.
- arXiv:2410.12628
- Balraj98. (2018). Stanford background dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/balraj98/stanford-background-dataset
- EKYC Solutions. (2022). Khmer OCR benchmark dataset (KHOB) [Data set]. GitHub. https://github.com/EKYCSolutions/khmer-ocr-benchmark-dataset
- Em, H., Valy, D., Gosselin, B., & Kong, P. (2024). Khmer text recognition dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/emhengly/khmer-text-recognition-dataset
