Team Ai
Apppublic

whosouravsharma/diffusiondb-sd15-lora

sourceHugging Facecreativeml-openrail-mupdated 10d agoView on Hugging Face
0likes
App README

DiffusionDB SD 1.5 LoRA

Generate images with Stable Diffusion 1.5 plus a rank-32 style LoRA fine-tuned on 13,598 prompt–image pairs from DiffusionDB. The adapter gives SD 1.5 a more saturated, higher-contrast, illustrative look, and it follows prompts just as well as the base model.

SD 1.5 base vs. with the LoRA, same prompt and seed

Set LoRA strength to 0 to render plain SD 1.5 at the same seed. At strength 0 the output is pixel-identical to the base model, so the slider is a true before/after comparison.

This Space records visitor IP addresses. See Privacy.

This Space is the frontend of a two-Space system. Image generation runs on a GPU in the `diffusiondb-sd15-lora-inference` Space.

How it works

  1. 1.You enter a prompt and, optionally, a negative prompt, step count, guidance, LoRA strength and seed.
  2. 2.This Space connects to the inference Space with gradio_client. If the inference Space is asleep, that connection wakes it (see Cold starts).
  3. 3.The inference Space loads SD 1.5 (fp16) and the LoRA adapter from the Hub, applies the requested LoRA strength, and generates one 512×512 image with the DPM-Solver++ scheduler.
  4. 4.The image and the seed used come back here and are shown.
Browser ──▶ UI Space (this, Gradio) ──gradio_client /generate──▶ Inference Space (T4 GPU)
                                                                   │  SD 1.5 fp16
                                                                   │  + LoRA checkpoint-4240
                                                                   │  DPM-Solver++ scheduler
            ◀──────────────── image + seed ────────────────────────┘

Components

ComponentRole
This SpaceGradio interface. It runs no model code, holds no weights and doesn't install torch
Inference SpaceLoads the pipeline and serves one /generate endpoint. Runs on a T4 GPU
Model repoLoRA checkpoints, training and evaluation code, and the full model card
Dataset repoThe cleaned training data, the cached latents, and the 50 held-out evaluation prompts

Keeping the interface separate means this Space doesn't need a GPU. Only the inference Space costs GPU time.

Model

Base`stable-diffusion-v1-5/stable-diffusion-v1-5`
AdapterLoRA rank 32 on the UNet attention layers, checkpoint-4240 (the final step)
Training10 epochs on 13,598 images at 512×512, with the text encoder frozen, in 2.7 hours on one A10G
Defaults here25 steps, guidance 7.5, LoRA strength 1.0 (adjustable from 0 to 1.5)

The example prompts in the UI are real prompts from the held-out evaluation set. The model never trained on them.

Evaluation at a glance

The model was evaluated on 506 held-out prompts, against the real DiffusionDB images for those prompts:

KID ×10³ ↓CLIP score
SD 1.5 base0.14 ± 0.3228.94
+ LoRA checkpoint-42402.88 ± 0.6128.98
  • —Prompt-following is unchanged. The two CLIP scores are within noise of each other.
  • —The LoRA adds its own style; it doesn't reproduce DiffusionDB's. DiffusionDB is itself SD 1.x output, and plain SD 1.5 already matches it (KID ≈ 0). The adapter moves outputs away from it, toward the saturated, high-contrast look in the gallery.
  • —No extra risk measured. The safety-checker flag rate is the same as base (2.6%), and a nearest-neighbour check finds no copying of training images.

The model card has the full protocol, the loss curve, the training progression, failure cases and every figure.

Intended use

Use it for giving SD 1.5 images a punchier, illustrative finish, and for seeing what a LoRA fine-tune changes compared with its base model.

Don't use it for muted or pastel palettes (the adapter overrides them), resolutions other than 512×512, or photorealistic images of real people.

API

The inference Space can be called directly:

python
from gradio_client import Client

client = Client("whosouravsharma/diffusiondb-sd15-lora-inference")
image_path, seed = client.predict(
    "a steampunk owl inside a glass jar, intricate detail",  # prompt
    "",      # negative prompt
    25,      # steps (at most 50)
    7.5,     # guidance
    1.0,     # LoRA strength; 0 = plain SD 1.5
    -1,      # seed; -1 = random
    api_name="/generate",
)

Cold starts

The inference Space goes to sleep after 5 minutes without requests. Waking it means scheduling a GPU, starting the container and loading about 2 GB of weights, so the first image after a wake is slow (the app estimates about 40 seconds). Later images take a few seconds.

This Space treats that as a wait, not an error. It retries the connection every 5 seconds for up to 10 minutes and shows Starting the server… with the time elapsed. It shows an error only when waiting can't help: when the inference Space is PAUSED, or in RUNTIME_ERROR, BUILD_ERROR or CONFIG_ERROR, it has to be restarted from its Settings page.

Privacy

For each generation, this Space records:

  • —the visitor's IP address, from the first entry of x-forwarded-for
  • —browser language, user agent and Gradio session ID
  • —the generation settings and the seed (never the image)

With LangSmith tracing on, which it is in this deployment, this is stored as run metadata in LangSmith. It is also written to the Space's container logs, which only the owner can see. An IP address counts as personal data under GDPR and UK GDPR, and under several US state laws.

VariableDefaultEffect
LANGSMITH_LOG_IPonoff stops recording the raw IP address
LANGSMITH_TRACINGunsettrue sends the data above to LangSmith

Setup (for forks / redeploys)

NameTypePurpose
HF_TOKENsecretRequired. Used to connect to the inference Space. Without it, every generation stops with an error
BACKEND_SPACEvariableOptional. Defaults to whosouravsharma/diffusiondb-sd15-lora-inference
LANGSMITH_API_KEYsecretOptional. Only needed with tracing
LANGSMITH_TRACING, LANGSMITH_PROJECTvariableOptional. Turn LangSmith tracing on and name the project

With tracing on, the UI and the inference Space appear in one trace, not two. The UI sends langsmith-trace and baggage headers with its request, so the backend's sd15_lora_generate run is recorded as a child of the UI's ui_generate run.

Limitations

  • —The style is fixed: saturation and contrast go up on every prompt, including ones that ask for muted colours.
  • —Some detailed scenes get simplified or lose their composition.
  • —Output is fixed at 512×512, the resolution the adapter was trained at.
  • —The training images were themselves made by SD 1.x, so the adapter reproduces that model's artifacts.
  • —The safety checker is off in the inference pipeline, so images are not screened. The training data was NSFW-filtered, but that filter is not a guarantee.
  • —The first image after the inference Space has been idle is slow (see Cold starts).

License

CreativeML OpenRAIL-M, the same license as the Stable Diffusion 1.5 base model. Training data comes from DiffusionDB (CC0 1.0; Wang et al., 2022, arXiv:2210.14896).