clearocr
Datasets
All datasets matching “clearocr”clearocr-invoice-document-ai
clearOCR Invoice Document AI Dataset
This dataset shows a complete invoice document AI workflow built around clearOCR.
It contains 423 high-confidence invoice examples with:
original invoice images,
OCR text generated by clearOCR,
Markdown reconstruction of the document,
structured invoice JSON generated by a local fine-tuned extraction model,
visual verification metadata.
The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.clearocr-benchmark-fiszki91-v1-results
OCR Bench Results: clearocr-benchmark-fiszki91-v1
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
clearocr.com/clearocr-api
1777
1734–1824
316
47
0
87%
2
lightonai/LightOnOCR-2-1B
1B
1567
1537–1603
222
141
0
61%
3
deepseek-ai/DeepSeek-OCR
4B
1412
1378–1444
137
225
0
38%
4
FireRedTeam/FireRed-OCR
2.1B
1380
1345–1411
120… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-benchmark-fiszki91-v1-results.clearocr-benchmark-pl-insurance-terms-v1
Document OCR using DeepSeek-OCR
This dataset contains markdown-formatted OCR results from images in byczong/pl-insurance-terms-struct using DeepSeek-OCR.
Processing Details
Source Dataset: byczong/pl-insurance-terms-struct
Model: deepseek-ai/DeepSeek-OCR
Number of Samples: 109
Processing Time: 14.0 min
Processing Date: 2026-03-28 19:57 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 8
Max Model Length: 8,192… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-benchmark-pl-insurance-terms-v1.clearocr-benchmark-pl-insurance-terms-v1-results
OCR Bench Results: Polish insurance terms and OWU benchmark
VLM-as-judge pairwise evaluation of OCR models on Polish insurance / OWU-style documents. Results depend strongly on document type, so this should be read as a document-specific benchmark rather than a universal OCR ranking.
This benchmark focuses on dense Polish insurance and legal text, including OWU-style documents with small fonts, long paragraphs, numbered sections and structured lists.
Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-benchmark-pl-insurance-terms-v1-results.clearocr-benchmark-fiszki91-v1
Document OCR using GLM-OCR
This dataset contains OCR results from images in Zombely/fiszki-ocr-test-A using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: Zombely/fiszki-ocr-test-A
Model: zai-org/GLM-OCR
Task: text recognition
Number of Samples: 91
Processing Time: 5.7 min
Processing Date: 2026-03-28 17:06 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-benchmark-fiszki91-v1.
