mansourkama/document-extraction-pipeline
0
Document Chat & Extraction Pipeline
Talk to your documents. Upload a PDF or image, then ask anything in plain language โ summarize, extract, restructure, translate, or export in any format. All models run free from HuggingFace, no API key required.
๐ [Live demo on HuggingFace Spaces](https://huggingface.co/spaces/mansourkama/document-extraction-pipeline)
What you can do
Plus template-based structured extraction (Invoice, Contract, Receipt, Personal Form) that returns a validated Pydantic model.
Architecture
Document (PDF or image)
โ
โผ
โโโโโโโโโโโโโโโโโ
โ OCR Engine โ doctr โ db_resnet50 (detection) + crnn_vgg16_bn (recognition)
โ (src/ocr.py) โ Pretrained weights from HuggingFace, no account needed
โโโโโโโโโฌโโโโโโโโ
โ raw text
โผ
โโโโโโโโโโโโโโโโโโโโโโโโ
โ LLM โ Qwen2.5-1.5B-Instruct (HuggingFace)
โ (src/extractor.py) โ โโ chat() โ free-form instruction โ any output format
โ โ โโ extract() โ schema-guided โ validated Pydantic model
โโโโโโโโโโโโโโโโโโโโโโโโProject Structure
document-extraction-pipeline/
โ
โโโ src/
โ โโโ ocr.py # OCR engine (doctr)
โ โโโ extractor.py # LLM: chat() + extract() methods
โ โโโ pipeline.py # Orchestrates OCR + LLM; exposes chat_file(), process()
โ โโโ templates/
โ โโโ base.py # ExtractionTemplate base class (schema_prompt + type coercion)
โ โโโ invoice.py # Invoice / billing
โ โโโ contract.py # Contract / mission order
โ โโโ receipt.py # Receipt / expense
โ โโโ form.py # Personal information form
โ
โโโ app.py # Gradio web interface (HuggingFace Spaces)
โโโ tests/
โ โโโ fixtures/generate_fixtures.py # Synthetic test images (Pillow)
โ โโโ test_pipeline.py # Unit + integration tests
โโโ examples/demo.py # CLI demo
โโโ requirements.txtQuickstart
pip install -r requirements.txtFree-form chat:
from src.pipeline import DocumentPipeline
pipeline = DocumentPipeline()
# From a file (OCR + LLM)
response = pipeline.chat_file("contract.pdf", "List all parties and the total value.")
print(response)
# From text (LLM only)
response = pipeline.chat(my_text, "Convert to a markdown table.")
print(response)Structured extraction:
from src.pipeline import DocumentPipeline
from src.templates import InvoiceTemplate
pipeline = DocumentPipeline()
result = pipeline.process("invoice.pdf", InvoiceTemplate)
print(result.model_dump_json(indent=2)){
"invoice_number": "INV-2024-0042",
"date": "2024-11-15",
"vendor_name": "TechSolutions SARL",
"total_amount": 8160.0,
"currency": "EUR"
}Built-in Templates
Custom template:
from pydantic import Field
from typing import Optional
from src.templates.base import ExtractionTemplate
class MedicalReportTemplate(ExtractionTemplate):
patient_name: Optional[str] = Field(None, description="Patient full name")
diagnosis: Optional[str] = Field(None, description="Main diagnosis")
physician: Optional[str] = Field(None, description="Treating physician name")
date: Optional[str] = Field(None, description="Report date")
result = pipeline.process("report.pdf", MedicalReportTemplate)Run Tests
python tests/fixtures/generate_fixtures.py # generate synthetic test images
pytest tests/ -vConfiguration
pipeline = DocumentPipeline(
ocr_det_arch="db_resnet50",
ocr_reco_arch="crnn_vgg16_bn",
llm_model="Qwen/Qwen2.5-1.5B-Instruct", # swap for any HF text-gen model
)