mrrobot2610/IDP-Machine-learning
IDP Machine Learning - Intelligent Document Processing
<div align="center">
Production-grade AI-powered document processing system for extracting structured data from documents
   
</div>
๐ฏ Overview
The IDP (Intelligent Document Processing) System is a complete end-to-end pipeline for:
- Document Classification - Automatically identifies document types (Invoice, Receipt, Form, Bank Statement)
- Named Entity Recognition - Extracts key fields (dates, amounts, IDs, names, addresses)
- OCR Integration - Text extraction from images and PDFs
Key Features
- ๐ Multi-format Support: PDF, PNG, JPEG, TIFF
- โก Fast Processing: <2 seconds per document on CPU
- ๐พ Lightweight: <500MB total memory footprint
- ๐ฏ High Accuracy: ~90% overall accuracy
๐ Model Performance
Accuracy Metrics
Performance Benchmarks
๐ง Models
1. Document Classifier
Supported Classes:
INVOICE- Invoices and billsRECEIPT- Purchase receiptsFORM- Application forms, tax formsBANK_STATEMENT- Bank statementsOTHER- Other document types
2. NER Model
Extracted Entities: | Entity | Example | |--------|---------| | INVOICE_NUMBER | INV-12345, #2024-001 | | DATE | 2024-01-15, Jan 15 2024 | | TOTAL_AMOUNT | $1,234.56, โน12,500.00 | | TAX_AMOUNT | $99.99, Tax: 18% | | VENDOR_NAME | Acme Corporation | | CUSTOMER_NAME | John Smith | | GST_ID | 27AAAC11234X1Z5 | | ADDRESS | 123 Main St, City |
3. OCR Engine
๐๏ธ Architecture
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ INFERENCE PIPELINE โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ 1. Preprocessing (OpenCV) โ
โ โโ Resize, Denoise, Deskew, Threshold, Enhance โ
โ โผ โ
โ 2. OCR (EasyOCR) โ
โ โโ Text + Bounding Boxes + Confidence โ
โ โผ โ
โ 3. Classification (MiniLM) โ
โ โโ Document Type + Confidence โ
โ โผ โ
โ 4. NER (DistilBERT) โ
โ โโ Entity Extraction (BIO tagging) โ
โ โผ โ
โ 5. Post-Processing โ
โ โโ Regex Fallbacks + Validation + Normalization โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ๐ Training Data
The models were trained on standard document understanding datasets:
Training Configuration
Classifier:
- Epochs: 15
- Batch Size: 16
- Learning Rate: 2e-5
- Early Stopping: 3 patience
NER:
- Epochs: 30
- Batch Size: 32
- Learning Rate: 3e-5
- Early Stopping: 5 patience
๐ Quick Start
Installation
# Clone the repository
git clone https://huggingface.co/mrrobot2610/IDP-Machine-learning
# Install dependencies
pip install -r requirements.txt
# System dependencies (PDF support)
# Ubuntu/Debian
sudo apt-get install poppler-utils
# macOS
brew install popplerUsage
from inference_pipeline import IDPPipeline
# Initialize pipeline
pipeline = IDPPipeline(
classifier_model_path="models/classifier/best_classifier.pt",
ner_model_path="models/ner/best_ner.pt",
use_gpu=False
)
# Process document
result = pipeline.process_document("invoice.pdf")
print(f"Document Type: {result['pages'][0]['document_type']}")
print(f"Fields: {result['pages'][0]['fields']}")API Server
# Start FastAPI server
python api_server.py
# Server runs on http://localhost:7860
# Health check
curl http://localhost:7860/health
# Process document
curl -X POST http://localhost:7860/process \
-F "file=@invoice.pdf"๐ Repository Structure
IDP-Machine-learning/
โโโ preprocessing.py # Image preprocessing (OpenCV)
โโโ ocr_engine.py # OCR integration (EasyOCR)
โโโ classifier_model.py # Document classifier model
โโโ ner_model.py # NER model for entity extraction
โโโ postprocessing.py # Output validation & formatting
โโโ inference_pipeline.py # Unified inference pipeline
โโโ api_server.py # FastAPI REST API
โโโ train_classifier.py # Classifier training script
โโโ train_ner.py # NER training script
โโโ dataset_loader.py # Dataset loading utilities
โโโ model_optimizer.py # ONNX conversion & quantization
โโโ demo_mode.py # Fallback rule-based logic
โโโ models/
โ โโโ classifier/
โ โ โโโ best_classifier.pt
โ โโโ ner/
โ โโโ best_ner.pt
โโโ frontend/ # Next.js frontend application
โโโ requirements.txt # Python dependencies๐ค API Response Format
{
"filename": "invoice.pdf",
"file_type": "pdf",
"total_pages": 1,
"pages": [{
"document_type": "INVOICE",
"classification_confidence": 0.96,
"fields": {
"invoice_number": {
"value": "INV-12345",
"confidence": 0.92,
"source": "ner"
},
"date": {
"value": "2024-01-15",
"confidence": 0.88,
"normalized": true
},
"total_amount": {
"value": "12500.00",
"numeric_value": 12500.0,
"currency": "INR",
"confidence": 0.95
}
},
"processing_time": {
"total": 0.92
}
}]
}๐ง Technology Stack
๐ Real-World Performance
Tips for Better Accuracy
- Use high-resolution scans (300+ DPI)
- Ensure good lighting for photos
- Enable adaptive thresholding for low-quality images
- Train on domain-specific data for best results
๐ค Contributing
Contributions are welcome! Please feel free to submit issues and pull requests.
๐ License
This project is licensed under the Apache 2.0 License.
๐ Acknowledgments
Built with:
Datasets
- CORD-v2 - Consolidated Receipt Dataset
- SROIE - ICDAR 2019 Competition Dataset
- FUNSD - Form Understanding Dataset
๐ง Contact
For questions and support, please open an issue in the repository.
