Team Ai
Modelpublic

Darayut/khmer-text-recognition

sourceHugging Facemitupdated 16d agoView on Hugging Face
2likes126downloads
Model Card

<div align="center"> <img src="assets/netra-logo-transparent-new.png" width="60%" alt="Netra Lab" /> </div>

<hr>

<p align="center"> <a href="https://github.com/netra-ai-lab/Netra-OCR"><b>GitHub</b></a> | <a href="https://pypi.org/project/netra-ocr/"><b>PyPI</b></a> | <a href="https://github.com/netra-ai-lab/Netra-OCR#netra-ocr-app-docker"><b>Desktop app (Docker)</b></a> | <a href="https://huggingface.co/collections/Darayut/khmer-text-synthetic"><b>Datasets</b></a> | <a href="https://huggingface.co/spaces/Darayut/Khmer-Text-Recognition"><b>Inference Space</b></a> </p>

<h2> <p align="center"> A Squeeze-and-Excitation Transformer Network for Khmer Optical Character Recognition </p> </h2>

<p align="center"> <img src="assets/benchmark.png" style="width: 1000px" align=center> </p>

<p align="center"> Character Error Rate (CER %) on KHOB, Legal Documents and Printed Words. <i>Lower is better.</i> </p>

Introduction

Netra-OCR is a 17M-parameter text-line recognizer for Khmer and English, trained from scratch on 1.3M synthetic text-line images. A Squeeze-and-Excitation network and a Transformer encoder read the image; a Transformer decoder writes the text. Khmer is decoded as grapheme clusters rather than single code points, which makes Khmer output sequences about 43% shorter than character-level decoding.

The model reads one text line at a time. To read whole pages (scans, phone photos, PDFs), use the `netra-ocr` package, which adds text-line detection, page layout analysis and export to Word, HTML, Markdown and JSON, or run the ready-made Netra OCR app with Docker (below).

Parameters17M
Inputone text-line image (any width; resized to 48 px high)
LanguagesKhmer, English
Output vocabulary2,101 tokens: Khmer grapheme clusters, Latin characters, digits, punctuation
Decodersar (autoregressive, default) and blockwise (identical output, ~1.55× faster)
Training data1.3M synthetic text lines
Packagepip install netra-ocr (version 1.1.0)

Files in this repository

FileWhat it is
khmerocr_cluster_ar.pthThe current model. Grapheme-cluster vocabulary, autoregressive decoder. The default for netra-ocr.
khmerocr_cluster_blockwise.pthThe same model with a trained blockwise-parallel decoding head. Produces exactly the same text as greedy ar, about 1.55× faster.
model.safetensors, config.json, *_khmerocr.py, inference.py, vocab.json, tokenizer filesAn earlier character-level model (124-token vocabulary), loadable with transformers + trust_remote_code=True. Kept for existing users; new projects should use the cluster model above.
khmerocr_epoch*.pthEarlier training checkpoints of the character-level model.

The cluster model's vocabulary (char2idx_cluster.json) ships inside the netra-ocr package, which downloads the checkpoint from this repository automatically.


Usage

Recognize text lines (Python)

bash
pip install netra-ocr
python
from netra_ocr.recognition import recognize, recognize_batch

# One cropped text line: a file path or a PIL.Image
text = recognize("line_crop.png")

# Many crops at once; the model stays loaded between calls
texts = recognize_batch(line_crops, batch_size=8)

# Blockwise decoder: same text, usually faster
text = recognize("line_crop.png", decoder="blockwise")

# Beam search: slower, slightly more accurate (autoregressive decoder only)
text = recognize("line_crop.png", beam_width=3)

The checkpoint is downloaded from this repository on first use and cached.

Read whole documents (Python)

bash
pip install "netra-ocr[yolo,layout,pdf,docx]"
python
from netra_ocr.ocr_engine import KhmerOCRPipeline
from netra_ocr.exporters import save_document

pipeline = KhmerOCRPipeline(detector="yolo", layout=True)
doc = pipeline.process_document("letter.pdf")     # images, multi-page TIFF or PDF
save_document(doc, "letter.docx")                 # .docx .html .md .json .txt

This detects the text lines, analyses the page layout (headings, paragraphs, tables, figures, headers and footers, reading order) with DocLayout-YOLO, and recognizes each line with this model.

Netra OCR app (no Python needed)

A web app and REST API with every model included. It runs on your own computer, on the CPU, with no internet connection needed after the download:

bash
docker run -p 8000:8000 -v netra-data:/data ghcr.io/netra-ai-lab/netra-ocr
# then open http://localhost:8000

Upload a scan, a phone photo or a PDF, correct the recognized document block by block next to the page image, and export it to Word. Works on Windows, Linux and Mac (Intel and Apple Silicon). See the GitHub README for details.


Datasets

The model was trained entirely on synthetic data and evaluated on real-world and synthetic data.

Training data (synthetic, 1.3M text lines)

DatasetLinesSourceAugmentations
khmer-document-synthetic-low-res100,000Pillow + Khmer corpus, 11 fontsErosion, noise, thinning/thickening, perspective distortion
khmer-scene-text-synthetic-contrast102,500SynthTIGER + Stanford BackgroundRotation, blur, noise, realistic backgrounds
KhmerSynthetic1M1,000,000SoyVitou (external)Pre-applied by the source authors
khmer-hanuman-100k100,000seanghay (external)Pre-applied by the source authors

Evaluation data

DatasetTypeSizeDescription
KHOBReal325Standard benchmark: clean backgrounds, compression artifacts.
Legal DocumentsReal227Smartphone photos of official documents (birth certificates, diplomas, ID cards): varied degradation, lighting and distortion. Not public.
Printed WordsSynthetic1,000Short, isolated words in 10 fonts, to test short sequences.

[image]


Methodology & Architecture

Preprocessing: chunking and merging

To handle variable-length text lines without aggressive resizing, each image is resized to a height of 48 px (keeping its aspect ratio) and split into overlapping 48×100 px chunks with a 16 px overlap.

Model architecture

[image]

<p><em>The input image is resized and split into 48×100 px chunks with 16 px overlaps. Each chunk passes through the Squeeze-and-Excitation network in parallel, producing 512 feature maps of 2×32 px, which become patch embeddings with positional embeddings. The Transformer encoder turns them into vision tokens; the tokens of all chunks are merged, smoothed by a bidirectional LSTM, and passed to the Transformer decoder, which outputs the text one character cluster at a time.</em></p>

  1. 1.Squeeze-and-Excitation network. Five VGG-style convolution blocks (64 → 512 channels) with height-only pooling in the later stages, so the horizontal axis, where character order lives, is preserved. Blocks 3, 4 and 5 add a 1D Squeeze-and-Excitation module that averages over height only, giving each horizontal column its own channel weights, so background and noise can be suppressed independently at each position along the line.

[image]

  1. 1.Patch module. Collapses the feature map's height and projects each column to a 384-dimensional embedding, plus a learnable positional embedding within the chunk.
  2. 2.Transformer encoder. Self-attention among the patches of a chunk resolves local ambiguities, such as visually similar sub-consonant stacks.
  3. 3.Merging module. Concatenates the vision tokens of all chunks of a line into one sequence and adds a second, global positional embedding across the whole line.
  4. 4.BiLSTM context smoother. A bidirectional LSTM runs over the merged sequence so information flows across chunk boundaries in both directions, smoothing the seams where a character is split between two chunks.

[image]

  1. 1.Transformer decoder. Generates the output sequence autoregressively, with cross-attention over the smoothed encoder output.

Tokenization

Latin text is decoded character by character. Khmer text is decoded as Khmer Character Clusters (a consonant with its subscripts and vowels, e.g. ក្រ), with any cluster outside the vocabulary falling back to single characters. On a 35,000-line Khmer corpus this shortens sequences by 42.65% with lossless round-tripping, which also means fewer decoding steps per line.

Blockwise-parallel decoding

khmerocr_cluster_blockwise.pth adds a small proposal head (Stern et al. 2018) on top of the frozen autoregressive model. Each step it proposes several tokens ahead and keeps only those that the base model itself would have produced, so the output is provably identical to greedy `ar` decoding, in fewer steps.

Training

Adam (learning rate 1e-4) with a staged cyclic learning-rate schedule, cross-entropy loss, and 50,000 randomly sampled, augmented lines per epoch. The cluster-vocabulary model was warm-started from the earlier character-level model (vision layers and shared token embeddings transferred), then trained on the full 1.3M-line set.


Evaluation

ModelKHOBLegal DocumentsPrinted Words
Tesseract5.141.88.0
Surya14.551.228.4
Qwen2.5-VL (3B) †89.284.698.6
DeepSeek-OCR (3B) †33.244.269.8
Netra-OCR (ours, 17M)1.05.33.0

Character Error Rate (%), lower is better. † Qwen2.5-VL and DeepSeek-OCR were fine-tuned with unsloth on the same training set as Netra-OCR before evaluation.

Decoders

The table above is the benchmark evaluation. This one compares the two decoders of the released checkpoints with each other, under identical settings: the 325 KHOB lines, greedy decoding, an NVIDIA RTX 3060.

DecoderCERWERExact matchms / line
ar (greedy)1.17%17.11%84.92%17.6
blockwise1.17%17.11%84.92%11.4

The blockwise decoder accepts on average 2.75 tokens per step (out of 4) and is 1.55× faster end to end, with identical output. Beam search (beam_width=3) lowers KHOB CER to about 0.9% at a higher cost per line.

Examples

<p align="center"> <img src="assets/eval1.png" style="width: 1000px" align=center> <img src="assets/eval2.png" style="width: 1000px" align=center> </p>


Limitations

  • —The model recognizes single text lines. Whole pages need a line detector first; the netra-ocr package includes one.
  • —It was trained on printed text; handwriting was not part of the training data.
  • —Only Khmer and English are supported.

License

The code and model weights are released under the MIT license. Note that part of the training data, KhmerSynthetic1M, is licensed by its authors for research and academic use only; check its terms before using this model commercially.


References

  1. 1.An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, et al. ICLR 2021. arXiv:2010.11929
  1. 1.TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei Florencio, Cha Zhang, Zhoujun Li, Furu Wei. AAAI 2023. arXiv:2109.10282
  1. 1.Toward a Low-Resource Non-Latin-Complete Baseline: An Exploration of Khmer Optical Character Recognition R. Buoy, M. Iwamura, S. Srun and K. Kise. IEEE Access, vol. 11, pp. 128044–128060, 2023. DOI: 10.1109/ACCESS.2023.3332361
  1. 1.Squeeze-and-Excitation Networks Jie Hu, Li Shen, and Gang Sun. CVPR 2018. arXiv:1709.01507
  1. 1.Bidirectional Recurrent Neural Networks Mike Schuster and Kuldip K. Paliwal. IEEE Transactions on Signal Processing, 1997. DOI: 10.1109/78.650093
  1. 1.Blockwise Parallel Decoding for Deep Autoregressive Models Mitchell Stern, Noam Shazeer, Jakob Uszkoreit. NeurIPS 2018. arXiv:1811.03115
  1. 1.DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception Zhiyuan Zhao, Hengrui Kang, Bin Wang, Conghui He.
  2. 2.arXiv:2410.12628
  1. 1.Balraj98. (2018). Stanford background dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/balraj98/stanford-background-dataset
  1. 1.EKYC Solutions. (2022). Khmer OCR benchmark dataset (KHOB) [Data set]. GitHub. https://github.com/EKYCSolutions/khmer-ocr-benchmark-dataset
  1. 1.Em, H., Valy, D., Gosselin, B., & Kong, P. (2024). Khmer text recognition dataset [Data set]. Kaggle. https://www.kaggle.com/datasets/emhengly/khmer-text-recognition-dataset