Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Voxel51 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/high-quality-invoice-images-for-ocr.image1K<n<10K6 likes3k downloads8mo agoHugging Face02drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face03typesafe /evalsafe-invoice-processing Invoice processing Snapshot: 2026-09-28. 150 cases and 6,874 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.tabular1K<n<10K6 likes1.3k downloads11d agoHugging Face04Francisco-Cruz /InvoicesReceiptsPTThis is a dataset comprising 1003 images of invoices and receipts, as well as the transcription of relevant fields for each document – seller name, seller address, seller tax identification, buyer tax identification, invoice date, invoice total amount, invoice tax amount, and document reference. It is organized as: folder 1_Images: files with pictures od the invoices/receipts folder 2_Annotations_Json: text files with the annotations on a json format Also available at:… See the full description on the dataset page: https://huggingface.co/datasets/Francisco-Cruz/InvoicesReceiptsPT.imagetext-classification1K<n<10K9 likes1.2k downloads2y agoHugging Face05JuanfelipeX123 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/JuanfelipeX123/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes816 downloads2mo agoHugging Face06katanaml-org /invoices-donut-data-v1 Dataset Card for Invoices (Sparrow) This dataset contains 500 invoice documents annotated and processed to be ready for Donut ML model fine-tuning. Annotation and data preparation task was done by Katana ML team. Sparrow - open-source data extraction solution by Katana ML. Original dataset info: Kozłowski, Marek; Weichbroth, Paweł (2021), “Samples of electronic invoices”, Mendeley Data, V2, doi: 10.17632/tnj49gpmtz.2 imagefeature-extractionn<1K44 likes659 downloads3y agoHugging Face07HV09 /synthetic-bilingual-invoices-200 Synthetic Bilingual Arabic/English Invoices — 200 documents with per-field ground truth 200 rendered invoice images and a matching 17-field ground-truth record for every one. Four language styles, 50 documents each: style what it exercises ar Arabic-only, Eastern-Arabic numerals (٤٤٬٥٤٨٫٣٥), RTL layout bilingual Arabic + English side by side, bidi field boundaries en English with Western numerals — the control en-au English (AU conventions) — different date/tax… See the full description on the dataset page: https://huggingface.co/datasets/HV09/synthetic-bilingual-invoices-200.imageimage-to-textn<1K0 likes560 downloads2mo agoHugging Face08mathieu1256 /FATURA2-invoicesThe dataset consists of 10000 jpg images with white backgrounds, 10000 jpg images with colored backgrounds (the same colors used in the paper) as well as 3x10000 json annotation files. The images are generated from 50 different templates. https://zenodo.org/records/10371464 dataset_info: features: - name: image dtype: image - name: ner_tags sequence: int64 - name: words sequence: string - name: bboxes sequence: sequence: int64 splits: - name: train… See the full description on the dataset page: https://huggingface.co/datasets/mathieu1256/FATURA2-invoices.imagefeature-extraction10K<n<100K21 likes510 downloads3y agoHugging Face09Lukaszl /clearocr-invoice-document-ai clearOCR Invoice Document AI Dataset This dataset shows a complete invoice document AI workflow built around clearOCR. It contains 423 high-confidence invoice examples with: original invoice images, OCR text generated by clearOCR, Markdown reconstruction of the document, structured invoice JSON generated by a local fine-tuned extraction model, visual verification metadata. The dataset demonstrates how clearOCR can serve as the OCR layer in an invoice automation pipeline where… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-invoice-document-ai.imageimage-to-textn<1K0 likes509 downloads5mo agoHugging Face10mychen76 /invoices-and-receipts_ocr_v1 Dataset Card for "invoices-and-receipts_ocr_v1" More Information needed image1K<n<10K90 likes419 downloads3y agoHugging Face11jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes415 downloads6mo agoHugging Face12osolmaz /invoice-fixture Invoice fixture Invoice Fixture is a synthetic dataset of financial documents, mostly in German, for testing document AI. Its 17 files are one invented person's paperwork from July 2026, the invoices and receipts and bank notices as they would pile up in a folder. Fourteen are PDFs, one of them a scan with no text layer, and three are JPEG images of the pages of drucker.pdf. Every name, address, account number and amount in the files was made up. They are test data and record… See the full description on the dataset page: https://huggingface.co/datasets/osolmaz/invoice-fixture.documentn<1K0 likes375 downloads4d agoHugging Face13chainyo /rvl-cdip-invoice⚠️ This only a subpart of the original dataset, containing only invoice. The RVL-CDIP (Ryerson Vision Lab Complex Document Information Processing) dataset consists of 400,000 grayscale images in 16 classes, with 25,000 images per class. There are 320,000 training images, 40,000 validation images, and 40,000 test images. The images are sized so their largest dimension does not exceed 1000 pixels. For questions and comments please contact Adam Harley (aharley@scs.ryerson.ca). The full dataset… See the full description on the dataset page: https://huggingface.co/datasets/chainyo/rvl-cdip-invoice.image10K<n<100K19 likes273 downloads5y agoHugging Face14Shubhal829 /high-quality-invoice-images-for-ocr Dataset Card for high_quality_invoice_images_ocr This is a FiftyOne dataset containing 8,181 high-quality synthetic invoice images for OCR and document understanding tasks. The dataset includes 1,489 fully annotated samples with structured JSON metadata and raw OCR text, plus 6,692 unannotated images for semi-supervised learning or annotation projects. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import… See the full description on the dataset page: https://huggingface.co/datasets/Shubhal829/high-quality-invoice-images-for-ocr.image1K<n<10K0 likes258 downloads4mo agoHugging Face15alamgirqazi /invoice-ocr-synthetic InvoiceOCR-Synth An annotation-noise-free synthetic dataset of receipt and invoice images for evaluating document information extraction systems, including vision–language models (VLMs) and OCR pipelines. DOI: 10.57967/hf/9733 Code: github.com/alamgirqazi/synthetic-invoice-gen Preprint: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7329421 Why this dataset Public receipt and invoice benchmarks rely on human annotation of pre-existing images. When an… See the full description on the dataset page: https://huggingface.co/datasets/alamgirqazi/invoice-ocr-synthetic.image1K<n<10K0 likes243 downloads23d agoHugging Face16deeptools-ai /test-document-invoiceimagen<1K2 likes205 downloads4y agoHugging Face17GokulRajaR /invoice-ocr-json Invoice OCR Dataset This dataset contains annotated invoice images and their corresponding OCR-extracted text in structured JSON format. The data was originally sourced from an open-source invoice dataset and processed using the GPT-4o mini model to extract relevant fields such as invoice number, date, total amount, vendor, and line items. Dataset Details Dataset Description This dataset is designed to support training and evaluation of document understanding… See the full description on the dataset page: https://huggingface.co/datasets/GokulRajaR/invoice-ocr-json.image1K<n<10K0 likes192 downloads1y agoHugging Face18mychen76 /invoices-and-receipts_ocr_v2 Dataset Card for "invoices-and-receipts_ocr_v2" Usage from datasets import load_dataset dataset = load_dataset("mychen76/invoices-and-receipts_ocr_v2") dataset More Information needed image1K<n<10K19 likes187 downloads2y agoHugging Face19Hemgg /invoices-and-receipts_ocr_v2image1K<n<10K0 likes162 downloads11mo agoHugging Face20Ananthu01 /7000_invoice_images_with_json_1The difference of this dataset with previous one is that all the JSON files contain the same keys and if the corresponding values are missing then they are given as null. Also, unlike the previous one, this dataset does not contain key-value pairs for everything in the corresponding invoice image, just only the important ones. The images are of size 448*448. image1K<n<10K0 likes161 downloads1y agoHugging Face21amaye15 /invoices-google-ocrimage10K<n<100K19 likes153 downloads2y agoHugging Face22Am0MuK /md_invoicesimagen<1K4 likes145 downloads3y agoHugging Face23HeliosMG4 /synthetic-english-invoices ENGLISH SYNTHETIC INVOICE DATASET — FREE SAMPLE PACK WHAT IS THIS? This is a free dataset of 20 100% synthetic (computer-generated) English invoices. No real person, company, or transaction is represented in these files. These are for AI/ML training and testing purposes only. WHAT IS INCLUDED? 20 invoice images (PNG format, high resolution) Standard English layout Fields: Company details, TRN, Buyer details, Item codes, Descriptions, Quantities, Unit prices, VAT (5%), and… See the full description on the dataset page: https://huggingface.co/datasets/HeliosMG4/synthetic-english-invoices.imagen<1K1 likes143 downloads23d agoHugging Face24raihan-js /invoice-check-jp Invoice-Check JP (synthetic) 4,100 synthetic Japanese qualified invoices (適格請求書) as images with gold JSON labels, plus the synthetic registry the invoices can be verified against. Built to measure how many wrong extractions symbolic checks (check digit, registry lookup, issuer-name match, tax arithmetic) catch, and how many slip through. Code and results: https://github.com/raihan-js/invoice-check-jp What is real and what is not Every company, address, phone… See the full description on the dataset page: https://huggingface.co/datasets/raihan-js/invoice-check-jp.imageimage-to-text1K<n<10K0 likes133 downloads5d agoHugging Face25ridwanFatur98 /synth-invoiceimage1K<n<10K0 likes132 downloads11d agoHugging Face26alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes129 downloads2mo agoHugging Face27KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes126 downloads4mo agoHugging Face28JohnTan38 /sparrow-invoice-v1textn<1K1 likes123 downloads3y agoHugging Face29laterrr /belege-de-invoices-sample1000 synthetic German invoices and credit notes (Rechnungen/Gutschriften), each a rendered image plus a JSON label with a pixel box for every field and every line item. This repository holds the free 40-document sample; the full 1000-document set is €14 (launch price) at j4zz.eu/belege. Everything below this paragraph, including the German DATASET.md text, is the same generator and the same fields, just a smaller draw for the sample. SROIE and FUNSD are the invoice/form datasets most… See the full description on the dataset page: https://huggingface.co/datasets/laterrr/belege-de-invoices-sample.object-detectionn<1K0 likes121 downloads1mo agoHugging Face30AlroWilde /invoice-checkmark-annotations Invoice Checkmark Annotations Multilingual dataset of real invoices with human-drawn visual checkmarks/circles indicating verified key fields. This dataset contains ~600 annotated invoice images (≈200 per language) in Ukrainian, Chinese, and Swedish. Each image shows real-world invoices where a human has manually added checkmarks (✓) or circles to highlight correctly extracted or verified fields (e.g. invoice number, buyer name, line totals, tax rate). Every sample includes:… See the full description on the dataset page: https://huggingface.co/datasets/AlroWilde/invoice-checkmark-annotations.0 likes112 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.