ocr
Datasets
All datasets matching “ocr”safedocs-cc-2m-paddle-vl-1-6-ocr
SafeDocs selected PaddleOCR-VL 1.6 OCR
OCR outputs for the PDFs accepted by the content-filtered selection. Each page row retains the complete native PaddleOCR result and its source document identity. Processing state is tracked in the run manifests.
OCRBenchGithub|Paper
OCRBench has been accepted by Science China Information Sciences.
Japanese-Political-Money-OCR-with-Qwenpost-ocr2short_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.Arxiv_2025_OCR
Arxiv_2025_OCR
OCR Data.
