Team Ai
Modelpublic

small-models-for-glam/kraken-ppocrv6-tiny

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
2likes
Model Card

PP-OCRv6 (tiny) multilingual text recognition base model

Hugging Face Hub mirror This repository mirrors the canonical Zenodo release by Benjamin Kiessling, ALMAnaCH, Inria Paris. The weights are byte-identical to the Zenodo release and have not been modified by Small Models for GLAM.

Use from the Hugging Face Hub

Install Kraken and the Hugging Face Hub CLI:

shell
pip install "kraken>=7.1.0" huggingface_hub

Download the model and run recognition:

shell
hf download small-models-for-glam/kraken-ppocrv6-tiny \
    tiny.safetensors \
    --local-dir ./kraken-ppocrv6-tiny

kraken -i image.png output.txt \
    segment -bl \
    ocr -m ./kraken-ppocrv6-tiny/tiny.safetensors

Description

This is the tiny variant (~0.69M parameters) of a family of PP-OCRv6 text-line recognition models (tiny, small, medium) for kraken. The models are trained from scratch with baseline

  • —bounding polygon data with a very diverse corpus containing historical, contemporary and born-digital document line images, handwritten and machine-printed, covering 44 languages across 10 scripts (Arabic, Armenian, Cyrillic, Ethiopic, Georgian, Greek, Hebrew, Latin, Malayalam, and Syriac).

Architecture

PP-OCRv6 is a conventional CTC-based line recognizer consisting of a lightweight convolutional backbone and a non-recurrent sequence-modelling neck.

The original architecture published as part of PaddlePaddle has been adapted for historical ATR by:

  • —increasing line height to 128px
  • —uncapping line width/CTC label budget
  • —replacing the optimizer with Adam+Muon

Uses

This is a base model that is supposed to produce usable output across a wide range of scripts and materials out of the box while also allowing fine-tuning with ease. It should offer similar accuracy and generalization to VLM-based recognizers without hallucinations and with vastly higher throughput. This medium variant model achieves the highest scores on the test set, tiny and small trade accuracy for inference speed and a smaller memory footprint.

Transcription guidelines, Normalization, and Transformations

No attempt has been made to normalize the source datasets to a single set of transcription guidelines; the corpus mixes conventions, so inconsistent output is to be expected, in particular for Latin-script European manuscripts which mix large datasets such as CATMuS and TRIDIS that have very different approaches to transcription. Text was normalized to Unicode NFD and whitespace was normalized during training and evaluation.

Bias, Risks, and Limitations

The training corpus is heavily skewed towards a handful of high-resource languages (English, French, German, Latin, Dutch, Middle French, ...). Languages with little real training material show markedly higher error rates and will require fine-tuning for practical use. Because transcription conventions are inconsistent across sources, the model may resolve abbreviations or expand glyphs unpredictably.

The synthetic data used for training was created with the pangoline tool which is limited to approximating modern, machine-printed text. For the languages/scripts present only as synthetic data (Classical Armenian, Geʽez) and to a lesser extent those sharing the Latin script (Irish, Latvian, Lithuanian, Romanian, Serbian, Slovenian), real-world accuracy is probably limited.

How to Get Started with the Model

Install kraken (>= 7.1.0), download the model, and run recognition on an input image:

shell
kraken -i image.png output.txt \
    segment -bl \
    ocr -m tiny.safetensors

For more information, refer to the documentation.

Training Details

Training Data

The model was trained on publicly available and restricted (private) datasets. Datasets marked as private are part of the training mixture but are not redistributable. A † marks languages that were additionally augmented with synthetic training data (see below).

LanguageScriptDatasets
Ancient Greek †GreekEPARCHOS, HPGTR, HTR_CPgr23, Stavronikita Monastery Greek Handwritten Document Collection No. 114, Stavronikita Monastery Greek Handwritten Document Collection No. 53, Stavronikita Monastery Greek Handwritten Document Collection No. 79, 11 private datasets
Arabic †ArabicAgapet, arabic_ms_data, iskandar, Muharaf: Manuscripts of Handwritten Arabic Dataset, OpenITI Arabic Print Data, RASAM dataset, TariMa
CatalanLatinFONDUE-CA-PRINT-20, htromance-spa, 1 private dataset
Church SlavonicCyrillic3 private datasets
Classical Armenian †Armeniansynthetic only
CorsicanLatinOCR Corse
Czech †Latin2024--medieval-czech-main, 2023--medieval-czech, HTR Winter School 2025 - Medieval Czech - Biblioteka Jagiellonska BJ Rkp 441 IV, Paderov Bible handwriting ground truth, ehri
DanishLatinehri
Dutch †Latin6000 ground truth of VOC and notarial deeds / HTR of VOC, WIC and notarial deeds, ARletta, Dagboek Ernest Clarysse, FONDUE-NE-MSS-17-PR
EnglishLatinFONDUE-EN-PRINT-20, IAM Handwriting Database, jcrs_train, jcrs_val, JosephHookerHTR, OCR-D gt_structure_text, sloanelab, The Revolutionary City / RevCity documentation, Memorials for Jane Lathrop Stanford, ehri, 2 private datasets
Finnish †LatinNewsEye/READ OCR Finnish Newspapers
FrenchLatinAntoine Verard extracts, Copiste-d-un-jour, corpus-HTR-lignes-mixtes, dataset-celestine-doniau-danest, FONDUE-FR-MSS-19, FONDUE-FR-MSS-19-PR, FONDUE-FR-PRINT-20, FONDUE-MLT-ART, genauto-td-htr, HTR Front Justice, La Correspondance Doucet-Rene Jean, Memoire sur St Domingue par H. M. Michel, Moonshines, NewsEye READ AS French Newspapers, NuBIS-OCR, Recensement Valaisan (Valais Time Machine), Tapus Corpus, TIMEUS Corpus, TitresNobiliaires_17_18, CREMMA Manuscrits du 20e, CREMMA Wikipedia, Maxime Kovalewsky - Coutume contemporaine et loi ancienne (1893), HTRomance, Modern Roman languages corpus, Peraire Ground Truth, PARES, 1 private dataset
Georgian †Georgian15 private datasets
GermanLatin2024--medieval-german, 2025--Early-Modern-German, Bullinger Digital Gwalther handwriting ground truth, charlottenburger-amtsschrifttum, Chronicling Germany, dach-gt, DigiTheo Ground Truth, Dresdner Hofdiarium, Fibeln, FONDUE-DE-MSS-16-PR, FONDUE-DE-MSS-18, FONDUE-DE-MSS-19-PR, FONDUE-DE-MSS-20-PR, FONDUE-MLT-ART, FONDUE-MLT-PRINT-TEST, FoNDUE_Kunsthistorisches-UZH_Archivdatenbank, Ground truth for Neue Zurcher Zeitung black letter, Hakenkreuzbanner, inzigkofen, Klosterneuburg, Stiftsbibl., Cod. 48, koenigsfelden, mkn-kurrent-gt, NewsEye / READ OCR Austrian Newspapers, nuremberg_letterbooks, OCR-D gt_structure_text, reichsanzeiger-gt, Training Data Incunabula Reichenau, Weisthuemer, gt-fraktur, german_kurrent_handwritten_text_lines, Ground Truth (Tagebücher Edwin Hennig), ehri, Fanny loves Wilhelm, Frauen im Fokus, Graphemic Early New German, 1 private dataset
German (shorthand)Latin1 private dataset
Geʽez †Ethiopicsynthetic only
HebrewHebrew2025-hebrew, 2 private datasets
HungarianLatinehri
Irish †Latinsynthetic only
ItalianLatinDiario del Soldato Bruno Celestino, EpiSearch HTR, FONDUE-IT-PRINT-20, FONDUE-IT-PRINT-20-PR, HTRogène Medieval Italian Manuscripts, HTRomance, Medieval Italian corpus of ground-truth for Handwritten Text Recognition, leopardi, LiDi1.0-project, LAM, 1 private dataset
LatinLatin2025--late-medieval-latin-main, burchards-dekret-digital, Caroline Minuscule ground truth, Carolingian Latin Group HTR Wien Winter School 2025, CREMMA Medii Aevi, DISTINGUO Latin ground truth, Eutyches, FONDUE-LA-MSS-16-PR, FONDUE-LA-MSS-17-PR, FONDUE-LA-MSS-MA, FONDUE-LA-PRINT-16, HTR Winter School 2023/2024 - Late Medieval Latin, ONB 3891, HTR Winter School 2024/2025 - Late Medieval Latin, ONB 4135; ONB 4680, HTRogène Medieval Latin Manuscripts, HTRomance, Medieval Latin corpus of ground-truth for Handwritten Text Recognition, notarial_charter, nubis, OCR-D gt_structure_text, Paris Bible Project, Training Data Incunabula Reichenau, Wien ONB Cod 2160 ground truth
Latvian †Latinsynthetic only
Lithuanian †Latinsynthetic only
MalayalamMalayalamGround Truth data for printed Malayalam
Middle DutchLatindata
Middle FrenchLatinCremma Medieval, De la généalogie des dieux, Données imprimés du 16e siècle, Données HTR incunables du 15e siècle, Données HTR manuscrits du 15e siècle, Données imprimés du 18e siècle, Données imprimés gothiques du 16e siècle, Fabliaux, FONDUE-FR-AAEB-16, FONDUE-FR-AAEB-17, FONDUE-FR-MSS-18, FONDUE-FR-PRINT-16, FONDUE-FR-PRINT-17, HTR-SETAF-Jean-Michel, HTR-SETAF-LesFaictzJCH, HTR-SETAF-Pierre-de-Vingle, HTRogene French, HTRomance, Medieval French corpus of ground-truth for Handwritten Text Recognition, Imprimés 17e siècle, Liber, OCR17plus, TNAH-2021-DecameronFR, transcription-chastel
Multilingual (mixed)LatinTraining Data Incunabula Reichenau, TranscriboQuest25_MedVernacReligio
NorwegianLatinNorHand v3 / Dataset for Handwritten Text Recognition in Norwegian
OccitanLatinHTRogène Medieval Occitan Manuscripts, 1 private dataset
Ottoman TurkishArabicmehmed_ibn_mehmed_uskubi_cukrikcizade_altiparmak_risale, OpenITI Arabic-script OCR Catalyst Project print/typeface data
Persian †Arabichafiz_divan, OpenITI Arabic-script OCR Catalyst Project print/typeface data, sadi_gulistan
PicardLatin1 private dataset
Polish †Latinehri
PortugueseLatiniForal-Dataset, Portuguese Handwriting 16th-19th c.
Romanian †Latinsynthetic only
Russian †Cyrillic2 private datasets
Serbian (Cyrillic) †Cyrillicsynthetic only
SlovakLatinSlovensky Supermodel P&T1, ehri
Slovenian †Latinsynthetic only
SpanishLatinFoNDUE Spanish chapbooks 19th c. Dataset, FONDUE-ES-MSS-19-PR, FONDUE-ES-PRINT-19, HTR - Araucania manuscript XIX, HTRogène Medieval Spanish Manuscripts, HTRomance, Medieval Spain corpus of ground-truth for Handwritten Text Recognition, ohg, 3 private datasets
SwedishLatinFinnish Court Records-sub500, kat57 Swedish ground truth dataset, NewsEye / READ OCR training dataset from Swedish Newspapers, riskarchiv
Syriac †Syriaczenodo.18157525, winter_school_vienna, 2 private datasets
UkrainianCyrillic1 private dataset
UrduArabicOpenITI Arabic-script OCR Catalyst Project print/typeface data
YiddishHebrew7 private datasets

Synthetic Training Data

Synthetic line images were generated as additional training material for 18 languages: Ancient Greek, Arabic, Classical Armenian, Czech, Dutch, Finnish, Georgian, Geʽez, Irish, Latvian, Lithuanian, Persian, Polish, Romanian, Russian, Serbian, Slovenian, Syriac. Of these, eight are present only as synthetic data: Classical Armenian, Geʽez, Irish, Latvian, Lithuanian, Romanian, Serbian (Cyrillic), Slovenian.

Training Procedure and Hyperparameters

The model was trained with kraken (feature/ppocrv6_rec branch).

Hardware3 × NVIDIA H100
Precisionbf16-mixed
OptimizerAdamW + Muon (momentum 0.95, weight decay 0.01)
Learning rate1.7e-3 (cosine schedule, 1500 step warmup, min 1e-6)
Batch size128
Gradient clipping1.0
Epochs16
NormalizationNFD + whitespace
Augmentationenabled

Evaluation

Testing Data

Metrics are computed on a held-out test set for each language. The test split was obtained by random 5% split with an upper limit of 100 document pages per language. No attempt has been made to split in a manner that separates documents between train and test. The scores below are therefore best read as in-domain generalization.

Evaluations with purely synthetic data are marked with ‡. CER and WER are the character- and word-level error rates, computed with torchmetrics using greedy CTC decoding and the NFD + whitespace normalization (equivalent to ketos test -u NFD -n).

Metrics

LanguageLinesCER (%)WER (%)
Ancient Greek45122.9887.79
Arabic1,92620.1168.77
Catalan1285.4726.59
Church Slavonic6,59917.8165.44
Classical Armenian ‡3,5800.643.23
Corsican501.8813.33
Czech92215.4761.28
Danish251.2910.89
Dutch4,38414.3151.81
English2,00514.4346.77
Finnish7,9071.166.83
French5,96714.0231.40
Georgian39428.7285.71
German3,0074.6117.96
German (shorthand)83041.2084.51
Geʽez ‡2,9921.275.58
Hebrew3,4269.1126.32
Hungarian383.2023.05
Irish ‡2,8241.004.15
Italian2,6117.2228.24
Latin5,74814.0344.34
Latvian ‡2,3971.256.48
Lithuanian ‡2,6151.306.73
Malayalam5948.7999.21
Middle Dutch3,01412.4443.63
Middle French3,9708.7032.38
Multilingual (mixed)2413.4917.99
Norwegian2,33514.9648.04
Ottoman Turkish45110.7048.25
Persian9909.3436.76
Polish2,7661.9310.21
Portuguese2,76323.5371.50
Romanian ‡2,4121.506.87
Russian3,05325.1168.33
Serbian (Cyrillic) ‡2,7890.582.71
Slovak1,5505.1121.50
Slovenian ‡2,7412.426.87
Spanish5,4656.1921.49
Swedish3,55511.8945.04
Syriac1,80110.7444.40
Ukrainian1,25313.7147.56
Urdu1,65610.2041.89
Yiddish4,3208.9632.79
Aggregate (micro-average)108,0108.7130.00
Aggregate (macro-average)10.9936.16

License

Released under the Apache-2.0 license.

Acknowledgements

Training of this model was funded by the European Union under Grant Agreement No.~101132163 (ATRIUM) and No.~101071829 (MiDRASH). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union.

This project also received funding from the BPI Scribe project.

Citation

If you use this model, please cite kraken and if possible credit the dataset providers linked in the front matter and the table above.