Team Ai
Modelpublic

ProCreations/Image-2.1-Turbo-Calibrated-FP8

sourceHugging Faceotherupdated 18h agoView on Hugging Face
1likes
Model Card

Image 2.1 Turbo Calibrated FP8

Built with Qwen. An independently calibrated W8A8 E4M3 quantization of the image transformer from Qwen-Image-2.1-Turbo (8-step distilled Qwen-Image-2.1), with an accelerated runtime. Research/evaluation only under the included Qwen Research License. This is not an official Qwen release.

This repository contains the quantized transformer (7.26 GB), an explicit loader and CLI, the full calibration/evaluation source, paired comparisons against BF16 Turbo, raw benchmark and profiler evidence, and a 30-second real-time demonstration. The BF16 text encoder, VAE, processor and the checkpoint's 8-step schedule load from the pinned upstream revision d65dbc9a7e8f6b5479e33dee6030eaab2a906509.

A sibling release, Image-2.1-Turbo-Calibrated-NVFP4, trades more fidelity for sub-second 1024² generation.

Speed

RTX PRO 6000 Blackwell (SM120, 96 GB), batch 1, all 8 saved-schedule steps, CFG 1, prefix KV cache.

Runtime1024 × 10242048 × 2048
BF16 Turbo (upstream Diffusers, eager)2.05 s11.45 s
FP8, eager1.53 s9.31 s
FP8, compiled, BF16 attention1.24 s7.63 s
FP8, release default (compiled + SageAttention2)1.17 s6.21 s

CUDA-synchronized end-to-end wall time: prompt encoding, all 8 transformer passes and VAE decode. Excluded: model load (about 4.3 s from local NVMe), first-use compilation in a fresh process (about 10 s at 1024, 15 s at 2048; once per resolution), and PNG writing. Five measured runs at 1024 and three at 2048 (five for the BF16-attention row), after one warmup. Peak GPU allocation: 28.6 GB at 1024 and 37.7 GB at 2048. Stage breakdown at 1024: encoder 16 ms, 8 transformer passes about 1.06 s, VAE 94 ms. Raw numbers: reports/benchmarks.

Quality against BF16 Turbo

18 held-out cases (16 text-to-image, 2 edits) that were not used for calibration, with matched seeds and resolutions (four at 2048², two non-square, the rest 1024²). Images are composited over white and resized to 512 px for LPIPS (AlexNet) and SSIM. Final latents are compared at full resolution.

RuntimeMean LPIPS ↓Mean SSIM ↑Mean PSNRLatent cosine
FP8, release default0.0550.94228.6 dB0.9891
FP8, eager0.0570.94528.1 dB0.9913

All 18 pairs were reviewed visually, along with all 25 images in the demo video. The English and Chinese titles stay correct, transparent RGBA output and both edits work, and hands, fur, feathers and glass look natural. No artifacts or general quality loss were observed. Outputs are not identical to BF16. Few-step models turn tiny numeric differences into composition changes: the worst pair (gouache harbour, LPIPS 0.18–0.24) keeps every element but draws the boats on sand instead of water. Compilation alone moves the eager FP8 output by mean LPIPS 0.020. SageAttention measured within run-to-run noise (0.043–0.057 over three paired runs, versus 0.056–0.060 with BF16 attention). See comparisons (BF16 | FP8 release | FP8 eager) and reports/quality_metrics.json. This is a finite fidelity check, not a human-preference study.

Precision and runtime

  • —224 projections (q/k/v/out and SwiGLU gate/proj/out in all 32 blocks) run as E4M3 W8A8: per-output-channel FP32 weight scales, dynamic per-token FP32 activation scales, per-channel FP32 SmoothQuant factors, native CUTLASS SM120 FP8 GEMM via torch._scaled_mm, FP32 accumulation (use_fast_accum=False), BF16 output. No layer needed a BF16 fallback. Conditioning, input/output projections, norms, the text encoder and the VAE stay BF16.
  • —Compiled decode: cached-prefix blocks run under torch.compile with dynamic shapes and emulated BF16 cast boundaries.
  • —Split prefill: step 1's prompt/reference prefix runs alone and fills the KV cache; the image tokens then take the same compiled path as later steps. Prefix tokens never attend to image tokens (causal_condition), so this is the same computation. FP8's per-token activation scaling makes the prefix KV cache bit-identical to upstream.
  • —VAE: single-frame decode skips Wan's unused video frame cache (bit-identical) and is compiled: 172 → 94 ms at 1024, 718 → 401 ms at 2048.
  • —Attention: --attention sage uses SageAttention2 (INT8 QK, FP8 PV, FP32 accumulation) in the cached passes, 1.39× faster than BF16 FlashAttention at 4096 tokens and 1.63× at 16384. --attention bf16 keeps native BF16 SDPA. The default auto uses Sage when installed.
  • —Not used: step skipping, feature caching, reduced resolution, distillation or LoRA.

Kernel evidence for one 1024² generation, from the profiler: 2,016 native SM120 FP8 GEMM launches, i.e. 224 × 9 transformer passes (step 1 prefix + target, then 7 cached steps); 256 SageAttention qk_int_sv_f8_attn_kernel launches (32 layers × 8 target passes); 32 masked memory-efficient SDPA launches for the prompt prefix; 36 BF16 FlashAttention launches in the 36-layer text encoder. One cached step takes 129 ms of GPU time (FP8 GEMMs 70%, SageAttention kernel 11%). See reports/kernel_evidence.json.

This is a custom Diffusers/PyTorch format that requires the included loader. It is not a drop-in Transformers, ComfyUI, vLLM or TensorRT checkpoint. Do not cast the loaded transformer to BF16/FP16. Tested only on SM120 with Torch 2.14.0+cu130, Diffusers 0.41.0, Transformers 5.17.0 and driver 615.71.09 (environment).

Calibration

  • —Trajectories: 64 deterministic BF16 Turbo trajectories on the checkpoint's 8-step schedule: 56 generations and 8 edits, 15 at 2048², 47 at 1024², and two 16:9/9:16. They cover photographs, portraits, textures, illustration, English/Chinese/Arabic/Japanese/French/Spanish prompts, typography, transparency and compositions.
  • —Activation rows: every projection samples 4 token rows at every step, 2,048 rows per layer: 512 fit rows and 512 disjoint diagnostic rows.
  • —Second moments: full-token input second moments E[xᵀx] come from all 3.62 M tokens of the same trajectories, using BF16 products with FP32 accumulation.
  • —Search per layer: five smoothing exponents × {round-to-nearest with three clipping factors, GPTQ error-feedback rounding}. Each candidate is scored with the actual FP8 kernel path.
  • —Result: GPTQ won for all 224 layers. Mean diagnostic output NRMSE is 1.77% (max 3.36%, mid-block attn.to_out.0). The best round-to-nearest candidate at the same smoothing averaged 2.18% fit NRMSE.

Everything is in calibration_search.json, quantization_summary.json and calibration_manifest.json. This is post-training calibration only, not QAT or distillation.

Run

Linux, Python 3.12, an NVIDIA Blackwell GPU and a CUDA 13 driver.

bash
hf download ProCreations/Image-2.1-Turbo-Calibrated-FP8 --local-dir Image-2.1-Turbo-Calibrated-FP8
cd Image-2.1-Turbo-Calibrated-FP8
python -m venv .venv && source .venv/bin/activate
pip install torch==2.14.0 torchvision==0.29.0 --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
# optional, faster attention (see requirements.txt): SageAttention 2.2.0 built from source
python generate.py --prompt "A kingfisher above a forest stream, wildlife photography" --width 1024 --height 1024 --output image.png

The 8-step schedule comes from the checkpoint, so there is no --steps. Defaults are 2048 × 2048 (upstream's native resolution class); use --width/--height, multiples of 16. Upstream aspect presets are 2048², 2400×1792, 1792×2400, 2528×1696, 1696×2528, 2752×1536 and 1536×2752. --warmup reports the first-use cost separately. --prompts-json list.json serves many prompts from one loaded pipeline. --cache-prompts reuses exact encoder outputs for repeated text-only prompts. --eager disables all acceleration. Pass --base /path/to/Qwen-Image-2.1-Turbo to use a local upstream snapshot; otherwise the pinned revision is fetched without its BF16 transformer.

Library use: pipe = load_pipeline(base, "transformer") from fp8_runtime.py, then accelerate_pipeline(pipe, attention="sage") from acceleration.py.

Transparent PNGs, editing and multiple references

bash
python generate.py --prompt "A cute mint-green baby dragon mascot, full body" --transparent --width 1024 --height 1024 --output dragon.png
python generate.py --image input.png --prompt "Change the jacket to blue, keep everything else" --width 1024 --height 1024 --output edited.png
python generate.py --image examples/ref-dragon-white.png examples/ref-whale-white.png \
  --prompt "Image 1 shows a mint-green baby dragon. Image 2 shows a blue cartoon whale. Create one new picture with the dragon standing next to the whale, both fully visible." \
  --width 1024 --height 1024 --output together.png

--transparent adds a transparency instruction and keeps the native alpha channel; no background removal is applied. The example dragon is 74% transparent pixels (examples/fp8-dragon.png). References keep command-line order (image 1, image 2, …). In our tests, transparent RGBA references combined with a short "place X next to Y" prompt lost the second character, and the BF16 Turbo pipeline did the same, so this is upstream behaviour. White-background references with an explicit description worked in BF16, FP8 and NVFP4 (examples/fp8-together-rgb.png). The text-only benchmark does not describe multi-reference speed.

Real-time video

demo/realtime-30s-fp8.mp4: 30 s, 1024², 30 fps, image only. One completed warmup image is shown at t = 0; after that, each change happens when an actual generation finishes, with all waits kept at 1× speed. The video shows 25 new images (mean 1.18 s each, release default runtime). Maximum capture lag is 1.1 ms, and every sampled decoded frame matches the image the capture log says was displayed (min PSNR 34.4 dB). See verification, timestamps and the contact sheet. The prompts are disjoint from calibration and evaluation.

Reproduce

source/ holds the experiment scripts in run order: verify_base.py, calibrate.py, hessians.py, quantize_fp8.py (--shard i/n workers, then one assembly run), evaluate.py, metrics.py, speedlab.py, benchmark.py, video.py, verify_video.py. common.py holds the workstation paths; change them for another installation. Calibration activations and second moments (about 30 GB) are regenerated locally and are not needed for inference. Unit checks: test_gptq.py (the GPTQ packer reproduces the reference quantizer bit-exactly when H = I), test_vae.py (bit-identical VAE) and test_split.py (split vs upstream prefill).

License and modifications

The original model is under the Qwen Research License: non-commercial research and evaluation, subject to its terms. Keep Notice and the license when redistributing. The modified weights carry a modification notice in their safetensors metadata and in transformer/quantization_config.json. ProCreations, 2026-10-09.