CuTIsolation/Z-Image-Turbo-W4A8
Z-Image-Turbo W4A8
W4A8 (4-bit weight, 8-bit activation) quantized weights for Z-Image-Turbo, made for ComfyUI using ComfyUI's native asym_w4a8_int8 quantized-diffusion format (Comfy Kitchen).
Both the diffusion model and the Qwen3-4B text encoder are quantized, so the whole pipeline fits in low VRAM.
Files
Original BF16 sizes: diffusion 12.3 GB, text encoder 8.0 GB.
Usage (ComfyUI)
Place the files in your ComfyUI models directory:
ComfyUI/
├── models/
│ ├── diffusion_models/
│ │ └── z_image_turbo_w4a8.safetensors
│ ├── text_encoders/
│ │ └── qwen_3_4b_w4a8.safetensors
│ └── vae/
│ └── flux1-vae.safetensorsThen use the standard Z-Image-Turbo text-to-image workflow with a Load Diffusion Model node pointed at z_image_turbo_w4a8.safetensors and a Load CLIP node pointed at qwen_3_4b_w4a8.safetensors.
Both files are detected automatically by ComfyUI (.comfy_quant metadata keys); no custom nodes are required. The text encoder must be loaded through the Qwen3-4B / Z-Image CLIP path (it does not need the pooled output).
Quality & Speed
Verified on an 8 GB VRAM GPU (RTX 4060 Laptop) at 1024x1024, 8 sampling steps:
Image quality is visually identical between BF16, W4A8 and the official int8_convrot checkpoint.
Quantization format
Per quantized Linear layer the file stores:
<key>.weight— int8, ConvRot-rotated packed int4[N, K/2]<key>.weight_s_rel— fp8 e4m3fn group scale[N, K/group_size]<key>.weight_s_channel— fp32 channel scale[N]<key>.weight_codebook— fp32 Lloyd-Max codebook[16]<key>.comfy_quant— uint8 JSON{"format": "asym_w4a8_int8", "group_size": 16, "convrot_groupsize": ...}
1D norms, biases, the embedding table and cap_embedder.1 are kept in BF16.
