InsecureErasure/Z-Image-Turbo-NVFP4
Z-Image Turbo - NVFP4 Mixed-Precision
Surgical mixed-precision quantization of Z-Image Turbo (6B S3-DiT), generated with `convert_to_quant`.
Formats: NVFP4 (baseline) + MXFP8 (sensitive layers) + BF16 (critical layers). Size: 4.84 GB (-58% vs BF16). Inference: ComfyUI + `comfy-kitchen`, Blackwell GPU (RTX 50xx / B100 / B200).
Also available: MXFP8 uniform quantization (6.23 GB, near-lossless).
- Prompt:
A bust portrait of a woman in her mid-twenties with messy dark hair tied in a loose bun, wearing a worn denim jacket over a gray hoodie.
She is leaning her elbows on a washing machine, her chin resting on her folded hands. Behind her, a row of industrial dryers against a tiled wall,
with one dryer door hanging open. Above the dryers, a handwritten sign taped to the wall says 'OUT OF ORDER' in black marker,
with a small smiley face drawn on it. To her left, a plastic basket overflows with unfolded clothes. To her right, a vending machine glows green,
displaying 'SOAP $1.50' on a small digital screen. The light is cool and buzzing, like fluorescent tubes overhead. She looks tired but amused
with a faint smirk.- Sampler/Scheduler: Euler/Simple
- Steps: 9
- CFG: 1.0
- Shift: 3.0
- Seed: 920698660737993
- Resolution: 1024 x 1536
Strategy
Uses per-layer sensitivity analysis via `quant_probe` and the DiT quantization literature (PTQ4DiT, ViDiT-Q, SemanticDialect, SVDQuant) to maximize quality-per-byte:
- ~190 tensors → NVFP4 (4-bit E2M1): baseline for most attention + FF weights
- ~100 tensors → MXFP8 (8-bit E4M3 + E8M0): attention outputs, gate projections (w1), mid-block adaLN
- ~20 tensors → BF16: last QKV, late adaLN modulations, refiner outputs
- ~110 tensors → BF16: norms, biases, embeddings (auto-excluded by
--zimage)
MXFP8-protected layers
BF16-protected layers
Refiner sub-graphs
Generation
#!/bin/bash
# NVFP4 baseline + MXFP8 for sensitive layers + BF16 at critical points.
# Refiners: block 0 fully MXFP8, block 1 outputs kept in BF16.
# Last QKV (layer 29), late adaLN (22-29), and refiner outputs in BF16.
# All main-trunk w1 (gate) projections in MXFP8.
convert_to_quant -i $1 \
--nvfp4 --zimage --comfy_quant --save-quant-metadata \
--custom-type mxfp8 \
--custom-layers "layers\.(10|16|26)\.attention\.qkv\.weight|layers\.(27|28)\.attention\.qkv\.weight|layers\.(0|1)\.attention\.out\.weight|layers\.(3|6|9|11|12|13|14|19|20|26)\.attention\.out\.weight|layers\.(27|28|29)\.attention\.out\.weight|layers\.(3|4|5|6|7|8|9|10|11|12|13|14|15|16|17|18|19|20|21|22|23|24|25|26)\.feed_forward\.w1\.weight|layers\.(27|28|29)\.feed_forward\.w1\.weight|layers\.(16|17|18|19|20|21)\.adaLN_modulation\.0\.weight|context_refiner\.(0|1)\.attention\.qkv\.weight|context_refiner\.(0|1)\.feed_forward\.w1\.weight|context_refiner\.(0|1)\.feed_forward\.w2\.weight|context_refiner\.(0|1)\.feed_forward\.w3\.weight|noise_refiner\.(0)\.attention\.(qkv|out)\.weight|noise_refiner\.(0)\.feed_forward\.(w1|w2)\.weight|noise_refiner\.(1)\.feed_forward\.w1\.weight" \
--exclude-layers "layers\.(29)\.attention\.qkv\.weight|layers\.(22|23|24|25|26)\.adaLN_modulation\.0\.weight|layers\.(27|28|29)\.adaLN_modulation\.0\.weight|context_refiner\.(0|1)\.attention\.out\.weight|context_refiner\.(1)\.feed_forward\.w2\.weight|noise_refiner\.(1)\.attention\.qkv\.weight|noise_refiner\.(1)\.attention\.out\.weight|noise_refiner\.(1)\.feed_forward\.w2\.weight|noise_refiner\.(0|1)\.feed_forward\.w3\.weight" \
--num-iter 6000 --top-p 0.35 --calib-samples 8192 \
--scale-optimization iterative --scale-refinement-rounds 2 \
--extract-lora --lora-rank 32 \
-o "${1%%.safetensors}-nvfp4.safetensors"Included files
Use the LoRA with variable strength in ComfyUI for improved fidelity.
Requirements
- Inference: CUDA 13.0+, PyTorch 2.10+, `comfy-kitchen`, Blackwell GPU (RTX 50xx / B100 / B200)
- Generation:
convert_to_quant >= 1.2.6,comfy-kitchen
Comparison
¹ Estimated on RTX 5060 (Blackwell) with comfy-kitchen CUDA kernels.
Methodology
Layer sensitivity was analyzed using `quant_probe`, which computes per-tensor excess kurtosis, dynamic range, and aspect ratio, then scores them against the model's own distribution to recommend *KEEP*, FP8, or NVFP4.
Recommendations were cross-referenced against the DiT quantization literature:
- PTQ4DiT (NeurIPS 2024) — salient channels in QKV + FFN, last blocks most affected
- ViDiT-Q (ICLR 2025) — metric-decoupled sensitivity: self-attention dominates visual quality
- HTG (2025) — channel-dependent outliers, severe in later blocks
- SemanticDialect (2026) — block-wise mixed-format validated for video DiTs
- SVDQuant (ICLR 2025) — low-rank branch absorbs 4-bit error, validated NVFP4
Credits
- Quantization engine: `convert_to_quant` by silveroxides
- Z-Image Turbo model by Tongyi-MAI
- ComfyUI integration via `comfy-kitchen`
- Layer sensitivity analysis via `quant_probe`
