Team Ai
Apppublic

Text-to-Document-Generation/PDF-Redaction-API

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes
STRUCTURE.md270 linesDownload Raw Back to root
1# Project Structure2 3```4pdf-redaction-api/5│6├── main.py                      # FastAPI application entry point7├── Dockerfile                   # Docker configuration for deployment8├── requirements.txt             # Python dependencies9├── README.md                    # Project documentation (for HuggingFace)10├── DEPLOYMENT.md               # Deployment guide11├── .gitignore                  # Git ignore rules12├── .dockerignore               # Docker ignore rules13│14├── app/                        # Application modules15│   ├── __init__.py            # Package initialization16│   └── redaction.py           # Core redaction logic (PDFRedactor class)17│18├── uploads/                    # Temporary upload directory19│   └── .gitkeep               # Keep directory in git20│21├── outputs/                    # Redacted PDF output directory22│   └── .gitkeep               # Keep directory in git23│24├── tests/                      # Test suite25│   ├── __init__.py26│   └── test_api.py            # API endpoint tests27│28└── client_example.py           # Example client for API usage29```30 31## File Descriptions32 33### Core Files34 35#### `main.py`36FastAPI application with endpoints:37- `POST /redact` - Upload and redact PDF38- `GET /download/{job_id}` - Download redacted PDF39- `GET /health` - Health check40- `GET /stats` - API statistics41- `DELETE /cleanup/{job_id}` - Manual cleanup42 43#### `app/redaction.py`44Core redaction logic:45- `PDFRedactor` class46- OCR processing with pytesseract47- NER using HuggingFace transformers48- Entity-to-box mapping49- PDF redaction with coordinate scaling50 51### Configuration Files52 53#### `requirements.txt`54Python dependencies:55- FastAPI & Uvicorn (API framework)56- Transformers & Torch (NER model)57- PyPDF (PDF manipulation)58- pdf2image (PDF to image conversion)59- pytesseract (OCR)60- Pillow (Image processing)61 62#### `Dockerfile`63Multi-stage build:641. Install system dependencies (tesseract, poppler)652. Install Python dependencies663. Copy application code674. Configure for port 7860 (HuggingFace default)68 69### Documentation70 71#### `README.md`72HuggingFace Space documentation:73- Features overview74- API endpoint documentation75- Usage examples (cURL, Python)76- Response format77- Local development setup78 79#### `DEPLOYMENT.md`80Step-by-step deployment guide:81- HuggingFace Spaces setup82- Git workflow83- Configuration options84- Security considerations85- Troubleshooting86- Cost estimation87 88### Testing & Examples89 90#### `tests/test_api.py`91Unit tests for API endpoints:92- Health check tests93- Upload validation tests94- Error handling tests95 96#### `client_example.py`97Example client implementation:98- Upload PDF99- Download redacted file100- Health check101- Statistics102 103## Data Flow104 105```106┌─────────────────────────────────────────────────────────┐107│ 1. Client uploads PDF                                   │108│    POST /redact with file                               │109└─────────────────────────────────────────────────────────┘110                          ↓111┌─────────────────────────────────────────────────────────┐112│ 2. FastAPI (main.py)                                    │113│    - Validates file                                     │114│    - Generates job_id                                   │115│    - Saves to uploads/                                  │116└─────────────────────────────────────────────────────────┘117                          ↓118┌─────────────────────────────────────────────────────────┐119│ 3. PDFRedactor (app/redaction.py)                       │120│    - perform_ocr() → Extract text + boxes               │121│    - run_ner() → Identify entities                      │122│    - map_entities_to_boxes() → Link entities to coords  │123│    - create_redacted_pdf() → Generate output            │124└─────────────────────────────────────────────────────────┘125                          ↓126┌─────────────────────────────────────────────────────────┐127│ 4. Response                                             │128│    - Return job_id and entity list                      │129│    - Save redacted PDF to outputs/                      │130└─────────────────────────────────────────────────────────┘131                          ↓132┌─────────────────────────────────────────────────────────┐133│ 5. Client downloads                                     │134│    GET /download/{job_id}                               │135└─────────────────────────────────────────────────────────┘136```137 138## Key Components139 140### 1. FastAPI Application (`main.py`)141 142**Endpoints:**143- RESTful API design144- File upload handling145- Background task cleanup146- CORS middleware for web access147 148**Features:**149- Automatic OpenAPI documentation at `/docs`150- JSON response models with Pydantic151- Error handling with HTTP exceptions152- Request validation153 154### 2. Redaction Engine (`app/redaction.py`)155 156**Pipeline Steps:**157 1581. **OCR Processing**159   - Convert PDF pages to images (pdf2image)160   - Extract text and bounding boxes (pytesseract)161   - Store image dimensions for coordinate scaling162 1632. **NER Processing**164   - Load HuggingFace model165   - Identify entities in text166   - Return entity types and character positions167 1683. **Mapping**169   - Create character span index for OCR words170   - Match NER entities to OCR bounding boxes171   - Handle partial word matches172 1734. **Redaction**174   - Scale OCR image coordinates to PDF points175   - Create black rectangle annotations176   - Write redacted PDF with pypdf177 178### 3. Docker Container179 180**Layers:**181- Base: Python 3.10 slim182- System packages: tesseract-ocr, poppler-utils183- Python packages: From requirements.txt184- Application code: Copied last for better caching185 186**Optimizations:**187- Multi-stage build (not used here, but possible)188- Minimal base image189- Cached dependency layers190- .dockerignore to reduce context size191 192## Environment Variables193 194Default configuration (can be overridden):195 196```bash197PYTHONUNBUFFERED=1        # Immediate log output198HF_HOME=/app/cache        # HuggingFace cache directory199```200 201## Port Configuration202 203- **Development**: 7860 (configurable in main.py)204- **Production (HF Spaces)**: 7860 (required)205 206## Directory Permissions207 208Ensure write permissions for:209- `uploads/` - Temporary PDF storage210- `outputs/` - Redacted PDF storage211- `cache/` - Model cache (created automatically)212 213## Adding New Features214 215### Add New Endpoint216 2171. Define in `main.py`:218```python219@app.get("/new-endpoint")220async def new_endpoint():221    return {"message": "Hello"}222```223 2242. Add response model if needed2253. Update README.md documentation2264. Add tests in `tests/test_api.py`227 228### Add New Redaction Option229 2301. Modify `PDFRedactor` class in `app/redaction.py`2312. Add parameter to `redact_document()` method2323. Update API endpoint in `main.py`2334. Document in README.md234 235### Add Authentication236 2371. Install: `pip install python-jose passlib`2382. Create `app/auth.py` with JWT logic2393. Add middleware to `main.py`2404. Protect endpoints with dependencies241 242## Best Practices243 2441. **Logging**: Use `logger` for all important events2452. **Error Handling**: Catch exceptions and return meaningful errors2463. **Validation**: Use Pydantic models for request/response validation2474. **Cleanup**: Always clean up temporary files2485. **Documentation**: Keep README.md and code comments updated2496. **Testing**: Add tests for new features250 251## Performance Considerations252 253### Bottlenecks2541. OCR processing (most time-consuming)2552. Model inference (NER)2563. File I/O257 258### Optimizations259- Lower DPI for faster OCR (trade-off with accuracy)260- Cache loaded models in memory261- Use async file operations262- Implement request queuing for high load263- Consider GPU for NER model264 265### Scaling266- Horizontal: Multiple container instances267- Vertical: Larger CPU/RAM allocation268- Caching: Redis for temporary results269- Queue: Celery for background processing270