Team Ai
Datasetpublic

himalaya-ai/ocr-document-processing-eval

ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.

sourceHugging Faceupdated 4mo agoView on Hugging Face
1likes678downloads
Dataset Card

ocrdocumentprocessing_eval

Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks.

  • —Repo: himalaya-ai/ocr-document-processing-eval
  • —Task: document_processing_ocr
  • —Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns.
  • —Optional fine-tuning/eval file: *.sharegpt.json with messages and images.

Core Columns

  • —id: unique sample identifier
  • —image: relative path to the image file
  • —ocr: ground-truth text label

Source Mix

  • —indic_vision_bench_deva_ocr ocr/test: 40%
  • —nayana_bench_deva_documents default/hi: 20%
  • —nayana_bench_deva_documents default/mr: 15%
  • —devanagari_digits_mixed default/train: 15%
  • —hindi_handwritten_word_ocr default/test: 10%

Notes

Generated by scripts/sample_ocr_eval_sets.py from the GLM fine-tuning workspace. If a source only exposes a train split, keep the deterministic held-out row ids out of SFT/training runs.