Team Ai
Modelpublic

BBuf/Qwen-Image-ModelOpt-FP8-SGLang

sourceHugging Faceotherupdated 6mo agoView on Hugging Face
0likes24downloads
Model Card

Qwen Image ModelOpt FP8 SGLang Transformer

This repository contains a SGLang-ready ModelOpt FP8 transformer override for `Qwen/Qwen-Image`. It only replaces the transformer weights; tokenizer, scheduler, VAE, and other non-transformer components are loaded from the original base model.

The checkpoint is intended for SGLang Diffusion with the Qwen Image FP8 support from sgl-project/sglang#23155.

Usage

bash
sglang generate \
  --backend=sglang \
  --model-id=Qwen-Image \
  --model-path Qwen/Qwen-Image \
  --transformer-path BBuf/Qwen-Image-ModelOpt-FP8-SGLang \
  --prompt "A futuristic cyberpunk city at night, neon lights reflecting on wet streets" \
  --width=1024 \
  --height=1024 \
  --num-inference-steps=50 \
  --guidance-scale=4.0 \
  --seed=42 \
  --num-gpus=1 \
  --dit-cpu-offload false \
  --dit-layerwise-offload false \
  --warmup \
  --save-output

H100 Validation Snapshot

Validation was run on one H100 GPU using rank0 with --backend=sglang. The FP8 image below is from the fixed checkpoint after keeping the validated sensitive Qwen Image fallback tensors in BF16.

Artifacts:

BF16, 1024x1024, 50 stepsFP8 fixed, 1024x1024, 50 steps
BF16 outputFP8 fixed output

Benchmark, warmup excluded:

MetricBF16FP8 fixedDeltaSpeedup
E2E latency13.589 s12.159 s-1.430 s (-10.5%)1.12x
Denoising stage12.929 s11.437 s-1.491 s (-11.5%)1.13x
Decoding stage58.55 ms52.30 ms-6.25 ms (-10.7%)1.12x
Text encoding599.85 ms666.43 ms+66.57 ms (+11.1%)0.90x

Notes:

  • —Validation prompt: A futuristic cyberpunk city at night, neon lights reflecting on wet streets.
  • —Validation settings: 1024x1024, 50 inference steps, guidance_scale=4.0, seed=42, --dit-cpu-offload false, --dit-layerwise-offload false, --warmup.
  • —Profiler artifacts were captured separately with profiler flags; those profiler timings include profiling overhead and are not used as benchmark latency numbers.

Conversion Notes

The checkpoint was converted from a NVIDIA ModelOpt FP8 export with SGLang's build_modelopt_fp8_transformer tool. Most linear weights are FP8. The validated fallback set keeps numerically sensitive tensors in BF16, including the Qwen Image image-MLP output projection family needed for normal image quality.