Text-to-Document-Generation/PDF-Redaction-API
0
1# Project Structure2 3```4pdf-redaction-api/5│6├── main.py # FastAPI application entry point7├── Dockerfile # Docker configuration for deployment8├── requirements.txt # Python dependencies9├── README.md # Project documentation (for HuggingFace)10├── DEPLOYMENT.md # Deployment guide11├── .gitignore # Git ignore rules12├── .dockerignore # Docker ignore rules13│14├── app/ # Application modules15│ ├── __init__.py # Package initialization16│ └── redaction.py # Core redaction logic (PDFRedactor class)17│18├── uploads/ # Temporary upload directory19│ └── .gitkeep # Keep directory in git20│21├── outputs/ # Redacted PDF output directory22│ └── .gitkeep # Keep directory in git23│24├── tests/ # Test suite25│ ├── __init__.py26│ └── test_api.py # API endpoint tests27│28└── client_example.py # Example client for API usage29```30 31## File Descriptions32 33### Core Files34 35#### `main.py`36FastAPI application with endpoints:37- `POST /redact` - Upload and redact PDF38- `GET /download/{job_id}` - Download redacted PDF39- `GET /health` - Health check40- `GET /stats` - API statistics41- `DELETE /cleanup/{job_id}` - Manual cleanup42 43#### `app/redaction.py`44Core redaction logic:45- `PDFRedactor` class46- OCR processing with pytesseract47- NER using HuggingFace transformers48- Entity-to-box mapping49- PDF redaction with coordinate scaling50 51### Configuration Files52 53#### `requirements.txt`54Python dependencies:55- FastAPI & Uvicorn (API framework)56- Transformers & Torch (NER model)57- PyPDF (PDF manipulation)58- pdf2image (PDF to image conversion)59- pytesseract (OCR)60- Pillow (Image processing)61 62#### `Dockerfile`63Multi-stage build:641. Install system dependencies (tesseract, poppler)652. Install Python dependencies663. Copy application code674. Configure for port 7860 (HuggingFace default)68 69### Documentation70 71#### `README.md`72HuggingFace Space documentation:73- Features overview74- API endpoint documentation75- Usage examples (cURL, Python)76- Response format77- Local development setup78 79#### `DEPLOYMENT.md`80Step-by-step deployment guide:81- HuggingFace Spaces setup82- Git workflow83- Configuration options84- Security considerations85- Troubleshooting86- Cost estimation87 88### Testing & Examples89 90#### `tests/test_api.py`91Unit tests for API endpoints:92- Health check tests93- Upload validation tests94- Error handling tests95 96#### `client_example.py`97Example client implementation:98- Upload PDF99- Download redacted file100- Health check101- Statistics102 103## Data Flow104 105```106┌─────────────────────────────────────────────────────────┐107│ 1. Client uploads PDF │108│ POST /redact with file │109└─────────────────────────────────────────────────────────┘110 ↓111┌─────────────────────────────────────────────────────────┐112│ 2. FastAPI (main.py) │113│ - Validates file │114│ - Generates job_id │115│ - Saves to uploads/ │116└─────────────────────────────────────────────────────────┘117 ↓118┌─────────────────────────────────────────────────────────┐119│ 3. PDFRedactor (app/redaction.py) │120│ - perform_ocr() → Extract text + boxes │121│ - run_ner() → Identify entities │122│ - map_entities_to_boxes() → Link entities to coords │123│ - create_redacted_pdf() → Generate output │124└─────────────────────────────────────────────────────────┘125 ↓126┌─────────────────────────────────────────────────────────┐127│ 4. Response │128│ - Return job_id and entity list │129│ - Save redacted PDF to outputs/ │130└─────────────────────────────────────────────────────────┘131 ↓132┌─────────────────────────────────────────────────────────┐133│ 5. Client downloads │134│ GET /download/{job_id} │135└─────────────────────────────────────────────────────────┘136```137 138## Key Components139 140### 1. FastAPI Application (`main.py`)141 142**Endpoints:**143- RESTful API design144- File upload handling145- Background task cleanup146- CORS middleware for web access147 148**Features:**149- Automatic OpenAPI documentation at `/docs`150- JSON response models with Pydantic151- Error handling with HTTP exceptions152- Request validation153 154### 2. Redaction Engine (`app/redaction.py`)155 156**Pipeline Steps:**157 1581. **OCR Processing**159 - Convert PDF pages to images (pdf2image)160 - Extract text and bounding boxes (pytesseract)161 - Store image dimensions for coordinate scaling162 1632. **NER Processing**164 - Load HuggingFace model165 - Identify entities in text166 - Return entity types and character positions167 1683. **Mapping**169 - Create character span index for OCR words170 - Match NER entities to OCR bounding boxes171 - Handle partial word matches172 1734. **Redaction**174 - Scale OCR image coordinates to PDF points175 - Create black rectangle annotations176 - Write redacted PDF with pypdf177 178### 3. Docker Container179 180**Layers:**181- Base: Python 3.10 slim182- System packages: tesseract-ocr, poppler-utils183- Python packages: From requirements.txt184- Application code: Copied last for better caching185 186**Optimizations:**187- Multi-stage build (not used here, but possible)188- Minimal base image189- Cached dependency layers190- .dockerignore to reduce context size191 192## Environment Variables193 194Default configuration (can be overridden):195 196```bash197PYTHONUNBUFFERED=1 # Immediate log output198HF_HOME=/app/cache # HuggingFace cache directory199```200 201## Port Configuration202 203- **Development**: 7860 (configurable in main.py)204- **Production (HF Spaces)**: 7860 (required)205 206## Directory Permissions207 208Ensure write permissions for:209- `uploads/` - Temporary PDF storage210- `outputs/` - Redacted PDF storage211- `cache/` - Model cache (created automatically)212 213## Adding New Features214 215### Add New Endpoint216 2171. Define in `main.py`:218```python219@app.get("/new-endpoint")220async def new_endpoint():221 return {"message": "Hello"}222```223 2242. Add response model if needed2253. Update README.md documentation2264. Add tests in `tests/test_api.py`227 228### Add New Redaction Option229 2301. Modify `PDFRedactor` class in `app/redaction.py`2312. Add parameter to `redact_document()` method2323. Update API endpoint in `main.py`2334. Document in README.md234 235### Add Authentication236 2371. Install: `pip install python-jose passlib`2382. Create `app/auth.py` with JWT logic2393. Add middleware to `main.py`2404. Protect endpoints with dependencies241 242## Best Practices243 2441. **Logging**: Use `logger` for all important events2452. **Error Handling**: Catch exceptions and return meaningful errors2463. **Validation**: Use Pydantic models for request/response validation2474. **Cleanup**: Always clean up temporary files2485. **Documentation**: Keep README.md and code comments updated2496. **Testing**: Add tests for new features250 251## Performance Considerations252 253### Bottlenecks2541. OCR processing (most time-consuming)2552. Model inference (NER)2563. File I/O257 258### Optimizations259- Lower DPI for faster OCR (trade-off with accuracy)260- Cache loaded models in memory261- Use async file operations262- Implement request queuing for high load263- Consider GPU for NER model264 265### Scaling266- Horizontal: Multiple container instances267- Vertical: Larger CPU/RAM allocation268- Caching: Redis for temporary results269- Queue: Celery for background processing270 