invoice
Datasets
All datasets matching “invoice”high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.InvoicesReceiptsPTThis is a dataset comprising 1003 images of invoices and receipts, as well as the transcription of relevant fields for each document – seller name, seller address, seller tax identification, buyer tax identification, invoice date, invoice total amount, invoice tax amount, and document reference.
It is organized as:
folder 1_Images: files with pictures od the invoices/receipts
folder 2_Annotations_Json: text files with the annotations on a json format
Also available at:… See the full description on the dataset page: https://huggingface.co/datasets/Francisco-Cruz/InvoicesReceiptsPT.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.evalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.invoices-donut-data-v1
Dataset Card for Invoices (Sparrow)
This dataset contains 500 invoice documents annotated and processed to be ready for Donut ML model fine-tuning.
Annotation and data preparation task was done by Katana ML team.
Sparrow - open-source data extraction solution by Katana ML.
Original dataset info: Kozłowski, Marek; Weichbroth, Paweł (2021), “Samples of electronic invoices”, Mendeley Data, V2, doi: 10.17632/tnj49gpmtz.2
invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.
