j4inam/reflections-encoder
0
Reflections SigLIP text-query encoder
Embeds a natural-language search query into a 768-d L2-normalized vector using the SigLIP2 text tower (onnx-community/siglip2-base-patch16-256-ONNX) — the same checkpoint the photo image embeddings use, so query and image vectors share one cosine space. The model is baked into the Docker image at build time and stays resident in RAM, so requests never pay a download or load cost.
How it fits in
Browser ──/gallery?q=…──▶ Vercel web app
│ POST /encode (Bearer $SIGLIP_ENCODER_TOKEN)
▼
this Space ──▶ {vector:[768]} (model resident in RAM)
│
Vercel cosine-ranks that vector against image embeddings in Neon (pgvector)
▼
ranked photos
Deploy: push to main (encoder/**) ─▶ GitHub Actions ─(OIDC, no stored token)─▶ hf upload ─▶ Space rebuilds
Warm: GitHub Actions cron ─every 6h─▶ GET /health (keeps the Space under HF's 48h sleep)API
POST /encode—{ "query": "..." }→{ "vector": [768 floats] }. RequiresAuthorization: Bearer $SIGLIP_ENCODER_TOKEN.GET /health— open liveness check for the keep-warm cron.
Operating
- Secret: set
SIGLIP_ENCODER_TOKENon the Space; it must match the web app's value (the web app sends it as the bearer token). - Deploy: automatic on push to
maintouchingencoder/**, via Trusted Publishing (OIDC) — no HF token is stored anywhere. - Model: downloaded once during the image build and baked in; container starts (including post-sleep wakes) load it from local disk, never the network.
