build-small-hackathon/kirana-invoice-train-data
Kirana Invoice Training Data — Indian FMCG Training dataset for the Kirana Detective project — an AI pipeline that audits distributor invoices for Indian kirana (grocery) stores. The repository contains two distinct sub-datasets used to fine-tune two separate models. Dataset Summary Sub-dataset Purpose Size Format synthetic_invoices/ OCR fine-tuning (MiniCPM-V) 500 images + annotations PNG + JSONL fmcg_catalog.json Product name normalization… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/kirana-invoice-train-data.
Kirana Invoice Training Data — Indian FMCG
Training dataset for the Kirana Detective project — an AI pipeline that audits distributor invoices for Indian kirana (grocery) stores. The repository contains two distinct sub-datasets used to fine-tune two separate models.
Dataset Summary
Both sub-datasets are fully synthetic — generated programmatically from a hand-curated SKU catalog. No real customer or business data is included.
Sub-dataset 1: Synthetic Invoice Images
Overview
500 invoice images rendered in pure Python (Pillow) across four realistic formats. Designed to teach MiniCPM-V to extract structured JSON from invoice photos regardless of format, quality, or layout.
Invoice Formats (125 images each)
Invoice Contents
Each invoice contains:
- 5–14 line items randomly sampled from the 200-SKU FMCG catalog
- Supplier (1 of 10 major Indian FMCG distributors) with GSTIN
- Buyer (1 of 8 kirana store archetypes) with GSTIN
- Invoice number in formats:
INV/2024-25/XXXXX,TAX/...,GST/...,BILL/... - Date between April 2024 and March 2025
- Pricing in range ₹8–₹480 per item (15% anomaly rate for training diversity)
- GST calculations at 0%, 5%, 12%, 18%, or 28% depending on product category
Suppliers represented:
HF Dataset Schema
The dataset on HuggingFace Hub is stored as Parquet with two columns:
The response string follows this structure:
{
"supplier": "Nestle India Ltd",
"gstin_supplier": "07AAACN0032R1ZX",
"buyer": "Ravi Provision Store",
"gstin_buyer": "33AABCR5678K1ZQ",
"invoice_number": "INV/2024-25/04821",
"date": "2024-09-14",
"line_items": [
{
"raw_name": "MAGGI NDL 70GM",
"quantity": 12,
"unit_price": 45.50,
"gst_rate": 18,
"total": 546.00
}
],
"subtotal": 3840.00,
"gst_total": 691.20,
"invoice_total": 4531.20
}Data Splits
Sub-dataset 2: FMCG Product Name Normalization Pairs
Overview
A structured catalog of 200 Indian FMCG SKUs with known abbreviations and aliases, used to generate 2,000 synthetic (raw, canonical) training pairs for MiniCPM5-1B.
Catalog Structure (fmcg_catalog.json)
Each entry:
{
"product_id": "maggi_masala_70g",
"canonical_name": "Nestle Maggi Masala Noodles 70g",
"hsn_code": "1902",
"gst_rate": 18,
"category": "noodles",
"brand": "Nestle",
"common_aliases": [
"MAGGI 70G",
"MAGGI NDL 70",
"MAGGI MSL 70G",
"MAGGI MASALA 70",
"MGI 70G",
"MAGGI 70GM"
]
}SKU Breakdown by Category
Augmentation Strategy
Each canonical SKU name is transformed into realistic raw invoice variants using rule-based augmentation:
Normalization Pair Format
Each training sample follows a chat template:
{
"messages": [
{
"role": "system",
"content": "You are an Indian FMCG product name normalizer. Given a raw product name from a distributor invoice, return ONLY the canonical product name. No explanation, no punctuation — just the canonical name."
},
{
"role": "user",
"content": "Invoice product name: \"MAGGI NDL 70GM\""
},
{
"role": "assistant",
"content": "Nestle Maggi Masala Noodles 70g"
}
]
}Data Splits
Downstream Models
This dataset is used to fine-tune two models in the Kirana Detective pipeline:
Known Limitations & Biases
Load from HuggingFace Hub
from datasets import load_dataset
import json
ds = load_dataset("build-small-hackathon/kirana-invoice-train-data")
sample = ds["train"][0]
image = sample["image"] # PIL Image — ready for model input
data = json.loads(sample["response"]) # parse the JSON string
print(data["supplier"])
print(data["line_items"])Citation
@misc{kirana_invoice_train_data_2026,
author = {Syed Naazim hussain},
title = {Kirana Invoice Training Data: Synthetic Indian FMCG Invoices for OCR and Product Normalization},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/datasets/build-small-hackathon/kirana-invoice-train-data}},
note = {Part of the Kirana Detective project}
}License
CC0 1.0 Universal (Public Domain Dedication) All data in this repository — synthetic invoice images, annotations, and the SKU catalog — is released to the public domain. No attribution required.
The generation scripts (generate_invoices.py, build_catalog.py) are licensed under MIT.
Version: 1.0 Last Updated: June 10, 2026
