Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face02typesafe /evalsafe-invoice-processing Invoice processing Snapshot: 2026-09-28. 150 cases and 6,874 question instances. Default reference: consensus. Labels are model-generated references. Data Load configuration cases, questions, or run_results; all have a test split. cases: one row per case_id, with the complete input in input_json, descriptive metadata_json, and openai, anthropic, and consensus labelsets. Decisions are grouped by policy_id and contain status, actions, and primary_action. questions:… See the full description on the dataset page: https://huggingface.co/datasets/typesafe/evalsafe-invoice-processing.tabular1K<n<10K6 likes1.3k downloads11d agoHugging Face03jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes415 downloads6mo agoHugging Face04alirezaaminzadeh /docflow-invoice-samples-fa DocFlow Invoice Samples — Persian & Bilingual Synthetic invoice dataset for evaluating DocFlow AI field extraction pipelines. Published by Aria AI Engineering Team. Dataset Summary Property Value Samples 50 (synthetic, OCR-friendly) Languages Persian (FA), English (EN) Formats PNG images + JSON annotations Use case Invoice OCR benchmarking, AP automation R&D Synthetic Yes — no real PII Fields Annotated vendor_name… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/docflow-invoice-samples-fa.imageimage-to-textn<1K0 likes129 downloads2mo agoHugging Face05cottonwood-development /synthetic-invoice-sample Invoice 500 — Free Sample (50 records) This is a free 50-record sample of the full 500-record commercial dataset. Every record in this sample is a real, valid invoice record — real records, just the data. you'd build by hand. What's in this sample 50 records in records.jsonl (NDJSON, one record per line) Schema-validated: every record parses against the real industry-standard schema before publishing Ready to use: import directly into any invoice pipeline — no… See the full description on the dataset page: https://huggingface.co/datasets/cottonwood-development/synthetic-invoice-sample.tabularothern<1K0 likes100 downloads4d agoHugging Face06necrasov-ilya /ru-invoice-extraction-benchmark Набор для извлечения данных из русскоязычных счетов 50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста. Разделы Раздел Документы Назначение development 30 разработка шаблонов и примеров validation 10 выбор настроек test 10 итоговая оценка Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.tabulartext-generationn<1K0 likes99 downloads29d agoHugging Face07processvenue /INVOICE_ANNOTATION_V2tabularimage-classification1K<n<10K0 likes67 downloads10mo agoHugging Face08Samarth-27 /carbon-mrv-invoice-emissions Synthetic Carbon MRV Invoice-to-Emissions Dataset A synthetic dataset that models the core pipeline used by carbon Measurement, Reporting & Verification (MRV) platforms: turning a business document line item (invoice, fuel receipt, electricity bill, freight charge) into a GHG Protocol Scope 1 / 2 / 3 classification and a calculated emissions value. It was built as reference dataset for learning and prototyping — specifically for training/evaluating models that do: Scope… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/carbon-mrv-invoice-emissions.tabulartabular-regression1K<n<10K0 likes35 downloads2mo agoHugging Face09processvenue /INVOICE_ANNOTATION_V1tabularimage-classification1K<n<10K0 likes27 downloads10mo agoHugging Face10albertosei /invoice-ner-amaye15-annotationstabular1K<n<10K0 likes15 downloads5mo agoHugging Face11Lemonator2707 /Invoice-Lam-SROIE-Conversationstabularn<1K0 likes14 downloads1y agoHugging Face12albertosei /invoice-ner-mychen76-annotationstabularn<1K0 likes11 downloads5mo agoHugging Face13AntonisD /Invoicestabularn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.