Team Ai
Modelpublic

TensorVizion/SDXL-Lightning-Quantized

sourceHugging Facecreativeml-openrail-mupdated 1mo agoView on Hugging Face
0likes906downloads
Model Card

SDXL Turbo – Quantized GGUF (Q2, Q4, Q8)

Quantized SDXL Turbo weights in GGUF format for fast, low‑VRAM image generation. This repo provides multiple quantization levels so you can pick the best trade‑off between speed, VRAM usage, and image quality.

Variants

  • —Q2_K – Ultra‑light, minimal VRAM, best for experimentation or very limited GPUs.
  • —Q4_K – Balanced speed/quality, recommended default for most users.
  • —Q8_0 – Near‑FP16 quality, higher VRAM, best for maximum fidelity.
VariantFormatApprox. SizeNotes
Q2\_KGGUF~1 GBLowest VRAM, fastest, lowest quality
Q4\_KGGUF~1.8–2 GBGood balance of quality and speed
Q8\_0GGUF~2.7 GBHighest quality, more VRAM needed

Tested on RTX 4060 8 GB and similar GPUs.


Usage (Python – llama-cpp-python style backends) pip install --upgrade llama-cpp-python

Example loading (adjust path and variant):

from llama_cpp import Llama

llm = Llama( modelpath="sdxlturboq4k.gguf", nctx=4096, ngpu_layers=-1, # offload as much as possible to GPU )


Inference Notes

  • —For 8 GB GPUs, Q4K and Q80 are both usable; Q2_K is ideal if you want to run other heavy apps in parallel.
  • —For 4–6 GB GPUs, Q2K or Q4K are recommended.
  • —Higher quantization (Q8_0) preserves more detail and coherence, but uses more VRAM and is slightly slower.

License

  • —Base model: SDXL Turbo under the CreativeML OpenRAIL-M license.
  • —By using these weights, you agree to the terms of the original SDXL/SDXL Turbo license and any downstream restrictions.

Acknowledgements

  • —Original SDXL Turbo model by Stability AI and contributors.
  • —Quantization and GGUF conversion by TensorVizion / thomas Barrie.