maheshsmc/D17-image-caption-decoder
0
Day 17 β Image Caption Generator (ViT + GPT2)
Track 2 (Coding) starter project. Minimal, productionβleaning skeleton with a clean Streamlit UI.
π Quickstart
# 1) Create env (optional)
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
# 2) Install minimal deps
pip install -r requirements.txt
# 3) Run UI
streamlit run app.pyBy default, the app uses a placeholder captioner (works without heavy deps). Toggle βUse real modelβ in the sidebar to run the actual ViT+GPT2 model (requires torch and transformers).
Real Model (optional)
pip install torch transformers
# Then in the UI, enable "Use real model"π§ Flow (Input β Output)
Image β preprocess (resize) β [ViT encoder] β embedding β [GPT2 decoder] β caption β sanitize β UIποΈ Project Layout
Day17_ImageCaption/
ββ app.py
ββ utils.py
ββ build_index.py
ββ requirements.txt
ββ README.md
ββ logs/
β ββ (inference_log.jsonl will appear after first run)
ββ docs/
ββ flow.txtπ Logging & Evaluation
- Inference runs are appended to
logs/inference_log.jsonl. - Convert to CSV for quick analysis:
python build_index.pyπ§° Hugging Face Spaces (Streamlit)
- Space SDK: Streamlit
- Runs on: CPU is fine (real model is slower). For better speed, choose GPU.
- Add a
requirements.txtwith the same deps; includetorchandtransformersif using real model. - Push this folder to a repo and connect to a Space.
π Notes
- This starter includes a light safety/sanitization step in
sanitize_caption. - For production, consider better filtering and evaluation harnesses.
Generated on 2025-09-09 06:35:49 (local build)
