datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pp-ocr-mnn-eval
PP-OCR MNN evaluation dataset (811-cell matrix)
Images (273) + canonical paddle.inference baselines (808 json) + configs
for scoring pp-ocr-mnn outputs. See README.md and
https://github.com/baicai1145/pp-ocr-mnn (tools/score.py).
A single-file snapshot is also included as ppocr-eval-dataset.tar.zst.
abl-ppocr-trainpp-ocrv6-smoke-v4
OCR with PP-OCRv6 TINY
Plain-text OCR results for images from davanstrien/ufo-ColPali, produced by
PaddlePaddle's PP-OCRv6
tiny pipeline (1.5M (0.4M det + 1.1M rec)).
Processing details
Source: davanstrien/ufo-ColPali
Model: PP-OCRv6_tiny (PP-OCRv6_tiny_det + PP-OCRv6_tiny_rec)
Tier: tiny (1.5M (0.4M det + 1.1M rec))
Recognition accuracy: 73.5%
Languages: 49 languages (en, zh only — no ja)
Engine: paddle_static
Samples: 3
Processing time: 0.41 min
Processing… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/pp-ocrv6-smoke-v4.PPOCR-weightsflux_one_word_ppocr_filteredppocrlabel
