Team Ai
Apppublic

whosouravsharma/diffusiondb-sd15-lora-inference

sourceHugging Facecreativeml-openrail-mupdated 2mo agoView on Hugging Face
0likes
App README

Inference backend

Generation backend for `diffusiondb-sd15-lora`. Loads Stable Diffusion 1.5 plus the LoRA adapter and exposes a single /generate endpoint. The minimal interface is for debugging; the real client is the UI Space.

This is the Space that needs a GPU. t4-small renders a 512×512 image in a few seconds; CPU takes minutes.

python
from gradio_client import Client

client = Client("whosouravsharma/diffusiondb-sd15-lora-inference")
image, seed = client.predict(
    "a steampunk owl inside a glass jar, intricate detail",  # prompt
    "",       # negative_prompt
    25,       # steps
    7.5,      # guidance
    1.0,      # lora_scale — 0 renders vanilla SD 1.5
    -1,       # seed, -1 for random
    api_name="/generate",
)

Files

  • —app.py — Gradio app exposing /generate
  • —src/pipeline.py — model loading and generation, no Gradio dependency
  • —src/infer.py — CLI: python3 -m src.infer "a prompt"
  • —src/tracing.py — optional LangSmith tracing + structured logging
  • —scripts/compare_checkpoints.py — contact sheet across checkpoints/scales

Observability

Every generation logs one structured line — seed, steps, guidance, lora_scale, device, latency, ms/step — with no configuration needed.

LangSmith tracing is opt-in. Unset, src/tracing.py no-ops and the package is never imported, so the serving path is unchanged:

bash
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY=<key>
export LANGSMITH_PROJECT=diffusiondb-sd15-lora-inference

On the Space, set LANGSMITH_API_KEY as a secret and the other two as plain variables.

Traces carry the generation parameters, the resolved device/dtype/checkpoint, and latency — but never the image itself: LangSmith stores run I/O as JSON, and a base64 512×512 PNG is ~400 KB per run, which gets truncated and makes the trace view unusable. Set LANGSMITH_LOG_IMAGES=1 to attach a ~10 KB WEBP thumbnail instead (LANGSMITH_THUMBNAIL_PX controls the size).

A generation that arrives from the UI Space is not a trace of its own. The UI sends langsmith-trace and baggage headers with its gradio_client call; app.py reads them off the Gradio request and continues that trace, so the sd15_lora_generate run here lands as a child of the UI's ui_generate run instead of in a separate tree. The visitor location the UI puts in baggage rides along, so it appears on this run too — the backend never sees the visitor directly, since the call reaches it server-to-server from the UI Space. A direct API call sends no such headers and simply starts its own trace.

Comparing checkpoints

Fixed seed, fixed prompts, every (checkpoint, lora_scale) combination tiled into one labelled PNG — the direct way to see what the adapter changed:

bash
python3 -m scripts.compare_checkpoints \
    --checkpoints base checkpoint-4240 \
    --scales 0.0 0.5 1.0 --seed 42 --out sweep.png

With tracing on, the sweep is one parent run with a child run per image.

Environment

VariableDefaultEffect
BASE_MODELstable-diffusion-v1-5/stable-diffusion-v1-5base checkpoint
LORA_REPOwhosouravsharma/diffusiondb-sd15-loraadapter repo
LORA_CHECKPOINTcheckpoint-4240base skips the adapter
ATTENTION_SLICINGautoauto = MPS only; true/off to force
LOG_LEVELINFOlog verbosity

Weights are pulled from the model repo on first request, not baked into the image, so switching checkpoints is a LORA_CHECKPOINT env change. The fp16 variant is used on GPU, so the cold-start download is ~2 GB rather than ~4.

Batch seeds: image i of a batch is generated with seed + i, so any single image is reproducible on its own at that seed with --count 1.