whosouravsharma/diffusiondb-sd15-lora-inference
Inference backend
Generation backend for `diffusiondb-sd15-lora`. Loads Stable Diffusion 1.5 plus the LoRA adapter and exposes a single /generate endpoint. The minimal interface is for debugging; the real client is the UI Space.
This is the Space that needs a GPU. t4-small renders a 512×512 image in a few seconds; CPU takes minutes.
from gradio_client import Client
client = Client("whosouravsharma/diffusiondb-sd15-lora-inference")
image, seed = client.predict(
"a steampunk owl inside a glass jar, intricate detail", # prompt
"", # negative_prompt
25, # steps
7.5, # guidance
1.0, # lora_scale — 0 renders vanilla SD 1.5
-1, # seed, -1 for random
api_name="/generate",
)Files
app.py— Gradio app exposing/generatesrc/pipeline.py— model loading and generation, no Gradio dependencysrc/infer.py— CLI:python3 -m src.infer "a prompt"src/tracing.py— optional LangSmith tracing + structured loggingscripts/compare_checkpoints.py— contact sheet across checkpoints/scales
Observability
Every generation logs one structured line — seed, steps, guidance, lora_scale, device, latency, ms/step — with no configuration needed.
LangSmith tracing is opt-in. Unset, src/tracing.py no-ops and the package is never imported, so the serving path is unchanged:
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=<key>
export LANGSMITH_PROJECT=diffusiondb-sd15-lora-inferenceOn the Space, set LANGSMITH_API_KEY as a secret and the other two as plain variables.
Traces carry the generation parameters, the resolved device/dtype/checkpoint, and latency — but never the image itself: LangSmith stores run I/O as JSON, and a base64 512×512 PNG is ~400 KB per run, which gets truncated and makes the trace view unusable. Set LANGSMITH_LOG_IMAGES=1 to attach a ~10 KB WEBP thumbnail instead (LANGSMITH_THUMBNAIL_PX controls the size).
A generation that arrives from the UI Space is not a trace of its own. The UI sends langsmith-trace and baggage headers with its gradio_client call; app.py reads them off the Gradio request and continues that trace, so the sd15_lora_generate run here lands as a child of the UI's ui_generate run instead of in a separate tree. The visitor location the UI puts in baggage rides along, so it appears on this run too — the backend never sees the visitor directly, since the call reaches it server-to-server from the UI Space. A direct API call sends no such headers and simply starts its own trace.
Comparing checkpoints
Fixed seed, fixed prompts, every (checkpoint, lora_scale) combination tiled into one labelled PNG — the direct way to see what the adapter changed:
python3 -m scripts.compare_checkpoints \
--checkpoints base checkpoint-4240 \
--scales 0.0 0.5 1.0 --seed 42 --out sweep.pngWith tracing on, the sweep is one parent run with a child run per image.
Environment
Weights are pulled from the model repo on first request, not baked into the image, so switching checkpoints is a LORA_CHECKPOINT env change. The fp16 variant is used on GPU, so the cold-start download is ~2 GB rather than ~4.
Batch seeds: image i of a batch is generated with seed + i, so any single image is reproducible on its own at that seed with --count 1.
