Team Ai
Modelpublic

mrrobot2610/IDP-Machine-learning

sourceHugging Faceapache-2.0updated 10mo agoView on Hugging Face
0likes
Model Card

IDP Machine Learning - Intelligent Document Processing

<div align="center">

Production-grade AI-powered document processing system for extracting structured data from documents

![License](https://opensource.org/licenses/Apache-2.0) ![Python](https://python.org) ![PyTorch](https://pytorch.org) ![Transformers](https://huggingface.co/transformers)

</div>

๐ŸŽฏ Overview

The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:

  • โ€”Document Classification - Automatically identifies document types (Invoice, Receipt, Form, Bank Statement)
  • โ€”Named Entity Recognition - Extracts key fields (dates, amounts, IDs, names, addresses)
  • โ€”OCR Integration - Text extraction from images and PDFs

Key Features

  • โ€”๐Ÿ“„ Multi-format Support: PDF, PNG, JPEG, TIFF
  • โ€”โšก Fast Processing: <2 seconds per document on CPU
  • โ€”๐Ÿ’พ Lightweight: <500MB total memory footprint
  • โ€”๐ŸŽฏ High Accuracy: ~90% overall accuracy

๐Ÿ“Š Model Performance

Accuracy Metrics

TaskMetricScore
Document ClassificationAccuracy92.3%
NER Field ExtractionF1 Score87.1%
Overall PipelineField Accuracy89.5%

Performance Benchmarks

MetricCPU (Intel i7)GPU (T4)
Single page processing1.2s0.4s
Memory usage450MB2.1GB
Throughput~50 docs/min~150 docs/min

๐Ÿง  Models

1. Document Classifier

PropertyValue
Base Modelnreimers/MiniLM-L6-H384-uncased
Parameters22M
TaskText Classification
Accuracy>90% on test set

Supported Classes:

  • โ€”INVOICE - Invoices and bills
  • โ€”RECEIPT - Purchase receipts
  • โ€”FORM - Application forms, tax forms
  • โ€”BANK_STATEMENT - Bank statements
  • โ€”OTHER - Other document types

2. NER Model

PropertyValue
Base Modeldistilbert-base-uncased
Parameters66M
TaskToken Classification (BIO tagging)
F1 Score>85% on test set

Extracted Entities: | Entity | Example | |--------|---------| | INVOICE_NUMBER | INV-12345, #2024-001 | | DATE | 2024-01-15, Jan 15 2024 | | TOTAL_AMOUNT | $1,234.56, โ‚น12,500.00 | | TAX_AMOUNT | $99.99, Tax: 18% | | VENDOR_NAME | Acme Corporation | | CUSTOMER_NAME | John Smith | | GST_ID | 27AAAC11234X1Z5 | | ADDRESS | 123 Main St, City |

3. OCR Engine

PropertyValue
EngineEasyOCR
Size~10MB
Speed<0.5s per page on CPU
LanguagesEnglish

๐Ÿ—๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                     INFERENCE PIPELINE                       โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚  1. Preprocessing (OpenCV)                                  โ”‚
โ”‚     โ””โ”€ Resize, Denoise, Deskew, Threshold, Enhance          โ”‚
โ”‚                              โ–ผ                              โ”‚
โ”‚  2. OCR (EasyOCR)                                           โ”‚
โ”‚     โ””โ”€ Text + Bounding Boxes + Confidence                   โ”‚
โ”‚                              โ–ผ                              โ”‚
โ”‚  3. Classification (MiniLM)                                 โ”‚
โ”‚     โ””โ”€ Document Type + Confidence                           โ”‚
โ”‚                              โ–ผ                              โ”‚
โ”‚  4. NER (DistilBERT)                                        โ”‚
โ”‚     โ””โ”€ Entity Extraction (BIO tagging)                      โ”‚
โ”‚                              โ–ผ                              โ”‚
โ”‚  5. Post-Processing                                         โ”‚
โ”‚     โ””โ”€ Regex Fallbacks + Validation + Normalization         โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ“š Training Data

The models were trained on standard document understanding datasets:

DatasetSizeDocument Type
CORD-v2~1,000 samplesReceipts
SROIE~1,000 samplesReceipts
FUNSD~200 samplesForms

Training Configuration

Classifier:

  • โ€”Epochs: 15
  • โ€”Batch Size: 16
  • โ€”Learning Rate: 2e-5
  • โ€”Early Stopping: 3 patience

NER:

  • โ€”Epochs: 30
  • โ€”Batch Size: 32
  • โ€”Learning Rate: 3e-5
  • โ€”Early Stopping: 5 patience

๐Ÿš€ Quick Start

Installation

bash
# Clone the repository
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning

# Install dependencies
pip install -r requirements.txt

# System dependencies (PDF support)
# Ubuntu/Debian
sudo apt-get install poppler-utils

# macOS
brew install poppler

Usage

python
from inference_pipeline import IDPPipeline

# Initialize pipeline
pipeline = IDPPipeline(
    classifier_model_path="models/classifier/best_classifier.pt",
    ner_model_path="models/ner/best_ner.pt",
    use_gpu=False
)

# Process document
result = pipeline.process_document("invoice.pdf")

print(f"Document Type: {result['pages'][0]['document_type']}")
print(f"Fields: {result['pages'][0]['fields']}")

API Server

bash
# Start FastAPI server
python api_server.py
# Server runs on http://localhost:7860

# Health check
curl http://localhost:7860/health

# Process document
curl -X POST http://localhost:7860/process \
  -F "file=@invoice.pdf"

๐Ÿ“ Repository Structure

IDP-Machine-learning/
โ”œโ”€โ”€ preprocessing.py          # Image preprocessing (OpenCV)
โ”œโ”€โ”€ ocr_engine.py             # OCR integration (EasyOCR)
โ”œโ”€โ”€ classifier_model.py       # Document classifier model
โ”œโ”€โ”€ ner_model.py              # NER model for entity extraction
โ”œโ”€โ”€ postprocessing.py         # Output validation & formatting
โ”œโ”€โ”€ inference_pipeline.py     # Unified inference pipeline
โ”œโ”€โ”€ api_server.py             # FastAPI REST API
โ”œโ”€โ”€ train_classifier.py       # Classifier training script
โ”œโ”€โ”€ train_ner.py              # NER training script
โ”œโ”€โ”€ dataset_loader.py         # Dataset loading utilities
โ”œโ”€โ”€ model_optimizer.py        # ONNX conversion & quantization
โ”œโ”€โ”€ demo_mode.py              # Fallback rule-based logic
โ”œโ”€โ”€ models/
โ”‚   โ”œโ”€โ”€ classifier/
โ”‚   โ”‚   โ””โ”€โ”€ best_classifier.pt
โ”‚   โ””โ”€โ”€ ner/
โ”‚       โ””โ”€โ”€ best_ner.pt
โ”œโ”€โ”€ frontend/                 # Next.js frontend application
โ””โ”€โ”€ requirements.txt          # Python dependencies

๐Ÿ“ค API Response Format

json
{
  "filename": "invoice.pdf",
  "file_type": "pdf",
  "total_pages": 1,
  "pages": [{
    "document_type": "INVOICE",
    "classification_confidence": 0.96,
    "fields": {
      "invoice_number": {
        "value": "INV-12345",
        "confidence": 0.92,
        "source": "ner"
      },
      "date": {
        "value": "2024-01-15",
        "confidence": 0.88,
        "normalized": true
      },
      "total_amount": {
        "value": "12500.00",
        "numeric_value": 12500.0,
        "currency": "INR",
        "confidence": 0.95
      }
    },
    "processing_time": {
      "total": 0.92
    }
  }]
}

๐Ÿ”ง Technology Stack

ComponentTechnology
Deep LearningPyTorch 2.x
NLP ModelsHugging Face Transformers
OCREasyOCR
Image ProcessingOpenCV
API FrameworkFastAPI + Uvicorn
FrontendNext.js 14 + React 18

๐Ÿ“ˆ Real-World Performance

Document QualityAccuracy
High-quality scans93-96%
Standard photos85-92%
Poor quality/handwritten65-80%

Tips for Better Accuracy

  • โ€”Use high-resolution scans (300+ DPI)
  • โ€”Ensure good lighting for photos
  • โ€”Enable adaptive thresholding for low-quality images
  • โ€”Train on domain-specific data for best results

๐Ÿค Contributing

Contributions are welcome! Please feel free to submit issues and pull requests.


๐Ÿ“„ License

This project is licensed under the Apache 2.0 License.


๐Ÿ™ Acknowledgments

Built with:

Datasets

  • โ€”CORD-v2 - Consolidated Receipt Dataset
  • โ€”SROIE - ICDAR 2019 Competition Dataset
  • โ€”FUNSD - Form Understanding Dataset

๐Ÿ“ง Contact

For questions and support, please open an issue in the repository.