datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Company-document-dataset-v2
Company Documents v2
Generation complete: all 13 document types have completed export and upload checkpoints.
Synthetic, born-digital business documents rendered from four open sample databases, with exact gold
labels: 353,580 PDFs (404,514 pages) of 13 document types in
English and French, issued by 60 synthetic companies,
each with its own letterhead, numbering and wording. Successor of
CompanyDocuments (2,677 PDFs, 4 types).
Dataset overview
property… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/Company-document-dataset-v2.clearocr-invoice-document-ai
clearOCR Invoice Document AI Dataset
This dataset shows a complete invoice document AI workflow built around clearOCR.
It contains 423 high-confidence invoice examples with:
original invoice images,
OCR text generated by clearOCR,
Markdown reconstruction of the document,
structured invoice JSON generated by a local fine-tuned extraction model,
visual verification metadata.
The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.
