datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.evalsafe-invoice-processing
Invoice processing
Snapshot: 2026-09-28. 150 cases and 6,874 question instances.
Default reference: consensus. Labels are model-generated references.
Data
Load configuration cases, questions, or run_results; all have a test split.
cases: one row per case_id, with the complete input in input_json, descriptive
metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped
by policy_id and contain status, actions, and primary_action.
questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.docflow-invoice-samples-fa
DocFlow Invoice Samples — Persian & Bilingual
Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines.
Published by Aria AI Engineering Team.
Dataset Summary
Property
Value
Samples
50 (synthetic, OCR-friendly)
Languages
Persian (FA), English (EN)
Formats
PNG images + JSON annotations
Use case
Invoice OCR benchmarking, AP automation R&D
Synthetic
Yes — no real PII
Fields Annotated
vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.synthetic-invoice-sample
Invoice 500 — Free Sample (50 records)
This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid invoice record — real records, just the data. you'd build by hand.
What's in this sample
50 records in records.jsonl (NDJSON, one record per line)
Schema-validated: every record parses against the real industry-standard schema before publishing
Ready to use: import directly into any invoice pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-invoice-sample.ru-invoice-extraction-benchmark
Набор для извлечения данных из русскоязычных счетов
50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста.
Разделы
Раздел
Документы
Назначение
development
30
разработка шаблонов и примеров
validation
10
выбор настроек
test
10
итоговая оценка
Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.INVOICE_ANNOTATION_V2carbon-mrv-invoice-emissions
Synthetic Carbon MRV Invoice-to-Emissions Dataset
A synthetic dataset that models the core pipeline used by carbon Measurement,
Reporting & Verification (MRV) platforms: turning a business document line
item (invoice, fuel receipt, electricity bill, freight charge) into a
GHG Protocol Scope 1 / 2 / 3 classification and a calculated emissions
value.
It was built as reference dataset for learning and prototyping —
specifically for training/evaluating models that do:
Scope… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/carbon-mrv-invoice-emissions.INVOICE_ANNOTATION_V1invoice-ner-amaye15-annotationsInvoice-Lam-SROIE-Conversationsinvoice-ner-mychen76-annotationsInvoices
