Team Ai
Apppublic

maheshsmc/D17-image-caption-decoder

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes
App README

Day 17 – Image Caption Generator (ViT + GPT2)

Track 2 (Coding) starter project. Minimal, production‑leaning skeleton with a clean Streamlit UI.

πŸš€ Quickstart

bash
# 1) Create env (optional)
python -m venv .venv && source .venv/bin/activate  # Windows: .venv\Scripts\activate

# 2) Install minimal deps
pip install -r requirements.txt

# 3) Run UI
streamlit run app.py

By default, the app uses a placeholder captioner (works without heavy deps). Toggle β€œUse real model” in the sidebar to run the actual ViT+GPT2 model (requires torch and transformers).

Real Model (optional)

bash
pip install torch transformers
# Then in the UI, enable "Use real model"

🧭 Flow (Input β†’ Output)

Image β†’ preprocess (resize) β†’ [ViT encoder] β†’ embedding β†’ [GPT2 decoder] β†’ caption β†’ sanitize β†’ UI

πŸ—‚οΈ Project Layout

Day17_ImageCaption/
β”œβ”€ app.py
β”œβ”€ utils.py
β”œβ”€ build_index.py
β”œβ”€ requirements.txt
β”œβ”€ README.md
β”œβ”€ logs/
β”‚  └─ (inference_log.jsonl will appear after first run)
└─ docs/
   └─ flow.txt

πŸ“’ Logging & Evaluation

  • β€”Inference runs are appended to logs/inference_log.jsonl.
  • β€”Convert to CSV for quick analysis:
bash
python build_index.py

🧰 Hugging Face Spaces (Streamlit)

  • β€”Space SDK: Streamlit
  • β€”Runs on: CPU is fine (real model is slower). For better speed, choose GPU.
  • β€”Add a requirements.txt with the same deps; include torch and transformers if using real model.
  • β€”Push this folder to a repo and connect to a Space.

πŸ”’ Notes

  • β€”This starter includes a light safety/sanitization step in sanitize_caption.
  • β€”For production, consider better filtering and evaluation harnesses.

Generated on 2025-09-09 06:35:49 (local build)