Umaim/image-captioning-seq2seq
1
1---2title: Image Captioning Seq2Seq3emoji: 🖼️4colorFrom: blue5colorTo: purple6sdk: docker7pinned: false8license: mit9---10 11# 🖼️ Image Caption Generator12 13An AI-powered image captioning application using ResNet-50 encoder and LSTM decoder with attention mechanism.14 15## Features16 17- **ResNet-50 Encoder**: Extracts visual features from images18- **LSTM Decoder**: Generates captions with attention mechanism19- **Two Decoding Strategies**:20 - Greedy Search: Fast, picks most probable word at each step21 - Beam Search: Explores multiple candidates for richer captions22- **Beautiful Dark Theme UI**: Modern, responsive interface23 24## Model Architecture25 261. **Image** → ResNet-50 (pretrained)272. **Features** → Linear Encoder (2048 → 512)283. **Decoder** → LSTM with Embedding (300d) + Attention294. **Output** → Natural language caption30 31## Training32 33- Dataset: Flickr8k34- Vocabulary Size: Dynamic (from training data)35- Max Caption Length: 20 tokens36 37## Usage38 391. Upload an image (JPG, PNG, or JPEG)402. Select decoding strategy (Greedy or Beam Search)413. Click "Generate Caption"424. View the AI-generated description43 44## Technical Stack45 46- **PyTorch** - Deep learning framework47- **Streamlit** - Web interface48- **torchvision** - Pre-trained ResNet-5049- **NLTK** - Text processing50 51## Files52 53- `app.py` - Main Streamlit application54- `caption_model.pkl` - Trained encoder/decoder weights + vocabulary55- `requirements.txt` - Python dependencies56- `Dockerfile` - Container configuration for deployment57 58## Local Development59 60```bash61# Install dependencies62pip install -r requirements.txt63 64# Run the app65streamlit run app.py66```67 68## Acknowledgments69 70Built with ❤️ using PyTorch and Streamlit