datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.evalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.InvoicesReceiptsPTThis is a dataset comprising 1003 images of invoices and receipts, as well as the transcription of relevant fields for each document – seller name, seller address, seller tax identification, buyer tax identification, invoice date, invoice total amount, invoice tax amount, and document reference.
It is organized as:
folder 1_Images: files with pictures od the invoices/receipts
folder 2_Annotations_Json: text files with the annotations on a json format
Also available at:… See the full description on the dataset page: https://huggingface.co/datasets/Francisco-Cruz/InvoicesReceiptsPT.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.invoices-donut-data-v1
Dataset Card for Invoices (Sparrow)
This dataset contains 500 invoice documents annotated and processed to be ready for Donut ML model fine-tuning.
Annotation and data preparation task was done by Katana ML team.
Sparrow - open-source data extraction solution by Katana ML.
Original dataset info: Kozłowski, Marek; Weichbroth, Paweł (2021), “Samples of electronic invoices”, Mendeley Data, V2, doi: 10.17632/tnj49gpmtz.2
synthetic-bilingual-invoices-200
Synthetic Bilingual Arabic/English Invoices — 200 documents with per-field ground truth
200 rendered invoice images and a matching 17-field ground-truth record for every
one. Four language styles, 50 documents each:
style
what it exercises
ar
Arabic-only, Eastern-Arabic numerals (٤٤٬٥٤٨٫٣٥), RTL layout
bilingual
Arabic + English side by side, bidi field boundaries
en
English with Western numerals — the control
en-au
English (AU conventions) — different date/tax… See the full description on the dataset page: https://huggingface.co/datasets/HV09/synthetic-bilingual-invoices-200.FATURA2-invoicesThe dataset consists of 10000 jpg images with white backgrounds, 10000 jpg images with colored backgrounds (the same colors used in the paper) as well as 3x10000 json annotation files. The images are generated from 50 different templates.
https://zenodo.org/records/10371464
dataset_info:
features:
- name: image
dtype: image
- name: ner_tags
sequence: int64
- name: words
sequence: string
- name: bboxes
sequence:
sequence: int64
splits:
- name: train… See the full description on the dataset page: https://huggingface.co/datasets/mathieu1256/FATURA2-invoices.clearocr-invoice-document-ai
clearOCR Invoice Document AI Dataset
This dataset shows a complete invoice document AI workflow built around clearOCR.
It contains 423 high-confidence invoice examples with:
original invoice images,
OCR text generated by clearOCR,
Markdown reconstruction of the document,
structured invoice JSON generated by a local fine-tuned extraction model,
visual verification metadata.
The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.invoices-and-receipts_ocr_v1
Dataset Card for "invoices-and-receipts_ocr_v1"
More Information needed
InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.invoice-fixture
Invoice fixture
Invoice Fixture is a synthetic dataset of financial documents, mostly in German, for testing
document AI. Its 17 files are one invented person's paperwork from July 2026, the invoices and
receipts and bank notices as they would pile up in a folder. Fourteen are PDFs, one of them a scan
with no text layer, and three are JPEG images of the pages of drucker.pdf.
Every name, address, account number and amount in the files was made up. They are test data and
record… See the full description on the dataset page: https://huggingface.co/datasets/osolmaz/invoice-fixture.rvl-cdip-invoice⚠️ This only a subpart of the original dataset, containing only invoice.
The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images. The images are sized so their largest dimension does not exceed 1000 pixels.
For questions and comments please contact Adam Harley (aharley@scs.ryerson.ca).
The full dataset… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/rvl-cdip-invoice.high-quality-invoice-images-for-ocr
Dataset Card for high_quality_invoice_images_ocr
This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.invoice-ocr-synthetic
InvoiceOCR-Synth
An annotation-noise-free synthetic dataset of receipt and invoice images for evaluating document information extraction systems, including vision–language models (VLMs) and OCR pipelines.
DOI: 10.57967/hf/9733
Code: github.com/alamgirqazi/synthetic-invoice-gen
Preprint: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7329421
Why this dataset
Public receipt and invoice benchmarks rely on human annotation of pre-existing images. When an… See the full description on the dataset page: https://huggingface.co/datasets/alamgirqazi/invoice-ocr-synthetic.test-document-invoiceinvoice-ocr-json
Invoice OCR Dataset
This dataset contains annotated invoice images and their corresponding OCR-extracted text in structured JSON format. The data was originally sourced from an open-source invoice dataset and processed using the GPT-4o mini model to extract relevant fields such as invoice number, date, total amount, vendor, and line items.
Dataset Details
Dataset Description
This dataset is designed to support training and evaluation of document understanding… See the full description on the dataset page: https://huggingface.co/datasets/GokulRajaR/invoice-ocr-json.invoices-and-receipts_ocr_v2
Dataset Card for "invoices-and-receipts_ocr_v2"
Usage
from datasets import load_dataset
dataset = load_dataset("mychen76/invoices-and-receipts_ocr_v2")
dataset
More Information needed
invoices-and-receipts_ocr_v27000_invoice_images_with_json_1The difference of this dataset with previous one is that all the JSON files contain the same keys and if the corresponding values are missing then they are given as null. Also, unlike the previous one, this dataset does not contain key-value pairs for everything in the corresponding invoice image, just only the important ones. The images are of size 448*448.
invoices-google-ocrmd_invoicessynthetic-english-invoices
ENGLISH SYNTHETIC INVOICE DATASET — FREE SAMPLE PACK
WHAT IS THIS?
This is a free dataset of 20 100% synthetic (computer-generated) English invoices.
No real person, company, or transaction is represented in these files.
These are for AI/ML training and testing purposes only.
WHAT IS INCLUDED?
20 invoice images (PNG format, high resolution)
Standard English layout
Fields: Company details, TRN, Buyer details, Item codes, Descriptions,
Quantities, Unit prices, VAT (5%), and… See the full description on the dataset page: https://huggingface.co/datasets/HeliosMG4/synthetic-english-invoices.invoice-check-jp
Invoice-Check JP (synthetic)
4,100 synthetic Japanese qualified invoices (適格請求書) as images with gold JSON labels, plus the synthetic registry the invoices can be verified against. Built to measure how many wrong extractions symbolic checks (check digit, registry lookup, issuer-name match, tax arithmetic) catch, and how many slip through. Code and results: https://github.com/raihan-js/invoice-check-jp
What is real and what is not
Every company, address, phone… See the full description on the dataset page: https://huggingface.co/datasets/raihan-js/invoice-check-jp.synth-invoicedocflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.arabic-latin-invoices-synthetic
Synthetic Arabic/Latin Invoices — label-first multimodal dataset
A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.sparrow-invoice-v1belege-de-invoices-sample1000 synthetic German invoices and credit notes (Rechnungen/Gutschriften),
each a rendered image plus a JSON label with a pixel box for every field and
every line item. This repository holds the free 40-document sample; the full
1000-document set is €14 (launch price) at j4zz.eu/belege.
Everything below this paragraph, including the German DATASET.md text, is
the same generator and the same fields, just a smaller draw for the sample.
SROIE and FUNSD are the invoice/form datasets most… See the full description on the dataset page: https://huggingface.co/datasets/laterrr/belege-de-invoices-sample.invoice-checkmark-annotations
Invoice Checkmark Annotations
Multilingual dataset of real invoices with human-drawn visual checkmarks/circles indicating verified key fields.
This dataset contains ~600 annotated invoice images (≈200 per language) in Ukrainian, Chinese, and Swedish. Each image shows real-world invoices where a human has manually added checkmarks (✓) or circles to highlight correctly extracted or verified fields (e.g. invoice number, buyer name, line totals, tax rate).
Every sample includes:… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/invoice-checkmark-annotations.
