bott-wa/multi-ai-document-analysis
0
Multi-AI Document Analysis ๐ค๐
Compare AWS Textract, Claude AI, and OpenAI GPT-4V for document processing and OCR tasks.
Features
๐ฅ 5 Processing Methods:
- AWS Textract (MCP) - Professional OCR with form/table detection
- Claude AI (Direct) - Vision + reasoning with Haiku model
- OpenAI GPT-4V (Direct) - Advanced vision + reasoning
- Hybrid (Textract + Claude) - OCR + AI enhancement
- Hybrid (Textract + OpenAI) - OCR + GPT-4V enhancement
๐ Cost Comparison:
- Budget: AWS Textract (~$0.0015/page)
- Balanced: Claude Hybrid (~$0.007/document)
- Premium: OpenAI Hybrid (~$0.015/document)
โจ Supported Formats: PNG, JPEG, PDF, TIFF (10MB max)
Setup
To run locally or fork this space:
- Add your API keys in the Space settings (or
.envfile locally):
AWS_ACCESS_KEY_ID=your_key
AWS_SECRET_ACCESS_KEY=your_secret
AWS_DEFAULT_REGION=us-east-1
ANTHROPIC_API_KEY=your_anthropic_key
OPENAI_API_KEY=your_openai_key- Install dependencies:
pip install -r requirements.txt- Run the app:
python app.pyUsage
- Test Connections: Verify your API keys work
- Upload Document: Drop any PDF, image, or form
- Select Method: Choose your processing approach
- Compare Results: See the differences in output and cost
Architecture
This app demonstrates a multi-AI comparison framework for document analysis:
- Direct API integrations (no cloud infrastructure needed)
- Structured JSON output for easy integration
- Cost and performance benchmarking
- Hybrid processing for enhanced accuracy
License
MIT License - feel free to fork and modify!
