datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InvoiceBenchmark
InvoiceBenchmark
200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number.
The Pitch
Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.INVOICE_ANNOTATION_V2carbon-mrv-invoice-emissions
Synthetic Carbon MRV Invoice-to-Emissions Dataset
A synthetic dataset that models the core pipeline used by carbon Measurement,
Reporting & Verification (MRV) platforms: turning a business document line
item (invoice, fuel receipt, electricity bill, freight charge) into a
GHG Protocol Scope 1 / 2 / 3 classification and a calculated emissions
value.
It was built as reference dataset for learning and prototyping —
specifically for training/evaluating models that do:
Scope… See the full description on the dataset page: https://huggingface.co/datasets/Samarth-27/carbon-mrv-invoice-emissions.INVOICE_ANNOTATION_V1Invoices
