Team Ai
Apppublic

mansourkama/document-extraction-pipeline

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes
App README

Document Chat & Extraction Pipeline

Talk to your documents. Upload a PDF or image, then ask anything in plain language โ€” summarize, extract, restructure, translate, or export in any format. All models run free from HuggingFace, no API key required.

๐Ÿš€ [Live demo on HuggingFace Spaces](https://huggingface.co/spaces/mansourkama/document-extraction-pipeline)


What you can do

Instruction exampleWhat you get
"Summarize in 3 bullet points"Concise summary
"Extract all dates and amounts as JSON"Structured JSON object
"List all parties and their roles"Named entity list
"Convert to a markdown table"Formatted table
"What are the payment terms?"Direct answer
"Translate the key fields to English"Translated output

Plus template-based structured extraction (Invoice, Contract, Receipt, Personal Form) that returns a validated Pydantic model.


Architecture

Document (PDF or image)
        โ”‚
        โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  OCR Engine   โ”‚   doctr โ€” db_resnet50 (detection) + crnn_vgg16_bn (recognition)
โ”‚  (src/ocr.py) โ”‚   Pretrained weights from HuggingFace, no account needed
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ”‚  raw text
        โ–ผ
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚  LLM                 โ”‚   Qwen2.5-1.5B-Instruct (HuggingFace)
โ”‚  (src/extractor.py)  โ”‚   โ”œโ”€ chat()    โ€” free-form instruction โ†’ any output format
โ”‚                      โ”‚   โ””โ”€ extract() โ€” schema-guided โ†’ validated Pydantic model
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Project Structure

document-extraction-pipeline/
โ”‚
โ”œโ”€โ”€ src/
โ”‚   โ”œโ”€โ”€ ocr.py                   # OCR engine (doctr)
โ”‚   โ”œโ”€โ”€ extractor.py             # LLM: chat() + extract() methods
โ”‚   โ”œโ”€โ”€ pipeline.py              # Orchestrates OCR + LLM; exposes chat_file(), process()
โ”‚   โ””โ”€โ”€ templates/
โ”‚       โ”œโ”€โ”€ base.py              # ExtractionTemplate base class (schema_prompt + type coercion)
โ”‚       โ”œโ”€โ”€ invoice.py           # Invoice / billing
โ”‚       โ”œโ”€โ”€ contract.py          # Contract / mission order
โ”‚       โ”œโ”€โ”€ receipt.py           # Receipt / expense
โ”‚       โ””โ”€โ”€ form.py              # Personal information form
โ”‚
โ”œโ”€โ”€ app.py                       # Gradio web interface (HuggingFace Spaces)
โ”œโ”€โ”€ tests/
โ”‚   โ”œโ”€โ”€ fixtures/generate_fixtures.py   # Synthetic test images (Pillow)
โ”‚   โ””โ”€โ”€ test_pipeline.py               # Unit + integration tests
โ”œโ”€โ”€ examples/demo.py             # CLI demo
โ””โ”€โ”€ requirements.txt

Quickstart

bash
pip install -r requirements.txt

Free-form chat:

python
from src.pipeline import DocumentPipeline

pipeline = DocumentPipeline()

# From a file (OCR + LLM)
response = pipeline.chat_file("contract.pdf", "List all parties and the total value.")
print(response)

# From text (LLM only)
response = pipeline.chat(my_text, "Convert to a markdown table.")
print(response)

Structured extraction:

python
from src.pipeline import DocumentPipeline
from src.templates import InvoiceTemplate

pipeline = DocumentPipeline()
result = pipeline.process("invoice.pdf", InvoiceTemplate)
print(result.model_dump_json(indent=2))
json
{
  "invoice_number": "INV-2024-0042",
  "date": "2024-11-15",
  "vendor_name": "TechSolutions SARL",
  "total_amount": 8160.0,
  "currency": "EUR"
}

Built-in Templates

TemplateDocument typeKey fields
InvoiceTemplateInvoices, billinginvoicenumber, date, vendor, client, subtotal, tax, total, currency, duedate
ContractTemplateContracts, mission ordersparties, missiondescription, startdate, enddate, totalvalue, payment_terms
ReceiptTemplateReceipts, expense slipsmerchant, date, total, currency, payment_method, items
PersonalFormTemplateRegistration, KYC formsfullname, dateof_birth, email, phone, address, nationality, occupation

Custom template:

python
from pydantic import Field
from typing import Optional
from src.templates.base import ExtractionTemplate

class MedicalReportTemplate(ExtractionTemplate):
    patient_name: Optional[str] = Field(None, description="Patient full name")
    diagnosis: Optional[str] = Field(None, description="Main diagnosis")
    physician: Optional[str] = Field(None, description="Treating physician name")
    date: Optional[str] = Field(None, description="Report date")

result = pipeline.process("report.pdf", MedicalReportTemplate)

Run Tests

bash
python tests/fixtures/generate_fixtures.py   # generate synthetic test images
pytest tests/ -v

Configuration

python
pipeline = DocumentPipeline(
    ocr_det_arch="db_resnet50",
    ocr_reco_arch="crnn_vgg16_bn",
    llm_model="Qwen/Qwen2.5-1.5B-Instruct",   # swap for any HF text-gen model
)

Stack

ComponentLibrary
Document OCRdoctr (HuggingFace pretrained)
Language modelQwen2.5-1.5B-Instruct
Schema validationPydantic v2
Web interfaceGradio ยท deployed on HuggingFace Spaces
Test fixturesPillow