Team Ai
Datasetpublic

osolmaz/invoice-fixture

Invoice fixture Invoice Fixture is a synthetic dataset of financial documents, mostly in German, for testing document AI. Its 17 files are one invented person's paperwork from July 2026, the invoices and receipts and bank notices as they would pile up in a folder. Fourteen are PDFs, one of them a scan with no text layer, and three are JPEG images of the pages of drucker.pdf. Every name, address, account number and amount in the files was made up. They are test data and record… See the full description on the dataset page: https://huggingface.co/datasets/osolmaz/invoice-fixture.

sourceHugging Faceupdated 5d agoView on Hugging Face
0likes375downloads
Dataset Card

[image]

Invoice fixture

Invoice Fixture is a synthetic dataset of financial documents, mostly in German, for testing document AI. Its 17 files are one invented person's paperwork from July 2026, the invoices and receipts and bank notices as they would pile up in a folder. Fourteen are PDFs, one of them a scan with no text layer, and three are JPEG images of the pages of drucker.pdf.

Every name, address, account number and amount in the files was made up. They are test data and record no real transaction.

Uses

The folder was built for testing PDF extraction and OCR, and for tasks where an agent has to work out which invoice a payment or a reimbursement belongs to. Some files repeat others on purpose. beispiel_2026_07.pdf and reimbursement/erstattung_gesamt.pdf each bundle documents that also exist on their own, so a reader that does not notice will count them twice.

Extraction benchmark

The `benchmark/` folder turns 14 of the documents into Harbor tasks. In each task an agent reads one document and writes its key fields to a JSON file, which a verifier then scores field by field against benchmark/ground_truth/. The task container has no OCR engine, so the agent has to read images with its own vision. Because the answers are public in this repository, the container can reach only the model's inference endpoint while the agent works.

In the round run on 2026-10-05 with Pi as the agent, GPT-6.1 Sol and GPT-6 Luna both reached a mean reward of 0.996, with 13 of the 14 documents fully correct. Ternary Bonsai 2 27B scored 0.992 on a GPU server, GLM-5.3-Flash 0.987 and DeepSeek-V4.1-Flash 0.920. The benchmark README lists every miss. The answer key covers only the fields of single documents and says nothing about how the documents relate to each other.

Files

fixtures/july-2026 holds the documents as separate files, with the reimbursement subfolder kept in its original nesting. manifest.jsonl lists the path, size, page count and SHA-256 of every file, and checksums.sha256 has the same checksums in the format that sha256sum -c reads.

Download

sh
hf download osolmaz/invoice-fixture --type dataset --local-dir invoice-fixture