small-models-for-glam/kraken-ppocrv6-tiny
PP-OCRv6 (tiny) multilingual text recognition base model
Hugging Face Hub mirror This repository mirrors the canonical Zenodo release by Benjamin Kiessling, ALMAnaCH, Inria Paris. The weights are byte-identical to the Zenodo release and have not been modified by Small Models for GLAM.
Use from the Hugging Face Hub
Install Kraken and the Hugging Face Hub CLI:
pip install "kraken>=7.1.0" huggingface_hubDownload the model and run recognition:
hf download small-models-for-glam/kraken-ppocrv6-tiny \
tiny.safetensors \
--local-dir ./kraken-ppocrv6-tiny
kraken -i image.png output.txt \
segment -bl \
ocr -m ./kraken-ppocrv6-tiny/tiny.safetensorsDescription
This is the tiny variant (~0.69M parameters) of a family of PP-OCRv6 text-line recognition models (tiny, small, medium) for kraken. The models are trained from scratch with baseline
- bounding polygon data with a very diverse corpus containing historical, contemporary and born-digital document line images, handwritten and machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian, Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).
Architecture
PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.
The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:
- increasing line height to 128px
- uncapping line width/CTC label budget
- replacing the optimizer with Adam+Muon
Uses
This is a base model that is supposed to produce usable output across a wide range of scripts and materials out of the box while also allowing fine-tuning with ease. It should offer similar accuracy and generalization to VLM-based recognizers without hallucinations and with vastly higher throughput. This medium variant model achieves the highest scores on the test set, tiny and small trade accuracy for inference speed and a smaller memory footprint.
Transcription guidelines, Normalization, and Transformations
No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.
Bias, Risks, and Limitations
The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.
The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.
How to Get Started with the Model
Install kraken (>= 7.1.0), download the model, and run recognition on an input image:
kraken -i image.png output.txt \
segment -bl \
ocr -m tiny.safetensorsFor more information, refer to the documentation.
Training Details
Training Data
The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A † marks languages that were additionally augmented with synthetic training data (see below).
Synthetic Training Data
Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.
Training Procedure and Hyperparameters
The model was trained with kraken (feature/ppocrv6_rec branch).
Evaluation
Testing Data
Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.
Evaluations with purely synthetic data are marked with ‡. CER and WER are the character- and word-level error rates, computed with torchmetrics using greedy CTC decoding and the NFD + whitespace normalization (equivalent to ketos test -u NFD -n).
Metrics
License
Released under the Apache-2.0 license.
Acknowledgements
Training of this model was funded by the European Union under Grant Agreement No.~101132163 (ATRIUM) and No.~101071829 (MiDRASH). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.
This project also received funding from the BPI Scribe project.
Citation
If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.
