build-small-hackathon/denali-toy-store
Denali Hang-Tag Grader
Snap it, we'll tag it. Photograph a returned toy and get a ready-to-shelf tag with its resale grade — powered by Rainier-VL-2B (~1.2B Zamba2-VL) on a patched llama.cpp, fully local: the model runs inside this Space, no cloud inference API in the path.
Space secrets: setHF_TOKEN(read access toDenali-AI/Rainier-VL-2B-Toys-GGUF) so the container can pull weights at startup.
Build Small Hackathon entry (Track 1 — Backyard AI). The submission is a Gradio app (hard rule) wearing a custom toy-store frontend (Off-Brand):
[ custom FX frontend / ]──→ POST /api/grade ─┐
[ Gradio surface /gradio ]─────────────────────┴→ 127.0.0.1:8001 (llama-server, OpenAI-compatible) → Rainier GGUFRun
# 1. start your model microservice on 127.0.0.1:8001 (patched llama.cpp)
# 2. start the app:
./run.sh # serves http://localhost:7860First-time setup: uv venv .venv --python 3.10 && uv pip install --python .venv/bin/python -r requirements.txt
No model running? Use the mock for development:
.venv/bin/python tools/mock_model_server.py # OpenAI-compatible stub on 127.0.0.1:8001URL dev hooks: ?sim (canned outcomes, no backend), ?demo&outcome=N (auto-runs a grade with a synthetic photo, deterministic), ?cam (jump to camera), append &live to force the real API in demo mode.
Test
.venv/bin/python -m pytest tests/ # 21 tests: grader rules, JSON parsing, API, e2e round-tripLayout
main.py— FastAPI host,/api/grade, Gradio app mounted at/gradiograder.py— verdict rules (reject if incomplete; discount if defect/creepy; else resell)inference.py— OpenAI-compatible client →http://127.0.0.1:8001/v1/chat/completions(prompts overridable viaINFERENCE_SYSTEM_PROMPT/INFERENCE_USER_PROMPT; base URL viaINFERENCE_BASE_URL)static/— the FX frontend, ported 1:1 from the Claude Design handoff in../hf-hackathon/(seeHACKATHON.md§3.3). The design is the spec — idle web render verified 100.00% pixel-identical to the prototype artboard.static/fonts/— self-hosted Fredoka / Nunito / JetBrains Mono (all SIL OFL, seeOFL-LICENSES.md); zero external requests at runtime.tools/— font fetcher, mock model servershots/— verification screenshots (idle/processing/result/error, web+mobile)
Design fidelity notes
Deliberate deviations from the prototype (per project decisions):
- Tweaks panel is design-tool chrome → defaults baked in (
config.py). - Mobile renders without the mock iOS status bar; safe-area padding instead.
- Desktop is a centered fluid canvas (artboard column metrics preserved); switches to the mobile layout ≤768px.
- "Printing your tag" keeps spinning until real inference answers, then the design's 420ms final beat plays — slow models read as intentional.
Social media post
Video included in LinkedIn post. https://www.linkedin.com/feed/update/urn:li:activity:7472401512586162176
