Team Ai
Modelpublic

wavespeed/Wan2.1-T2V-14B-Diffusers-fp16-nf4

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes4downloads
Model Card

Wan2.1-T2V-14B-Diffusers-fp16-nf4

NF4 weight-only quantized Wan2.1-T2V-14B-Diffusers, saved as a complete WanPipeline.

What was changed

All 40 blocks of the WanTransformer3DModel are quantized to nf4_weightonly — 4-bit NormalFloat weights with a rank-32 low-rank error correction term fitted over 8 iterations, which recovers most of the quality NF4 alone gives up. Activations stay in fp16.

The UMT5 text encoder is not quantized (it is kept in fp16); the VAE is loaded in bf16 and the scheduler is swapped to UniPCMultistepScheduler. Quantization was done with WaveSpeed's xelerate.ao.quantize. Weights are pickled .bin shards, so loading requires use_safetensors=False.

Usage

python
import torch
from diffusers import WanPipeline

pipe = WanPipeline.from_pretrained(
    "wavespeed/Wan2.1-T2V-14B-Diffusers-fp16-nf4",
    torch_dtype=torch.float16,
    use_safetensors=False,
).to("cuda")

Running the NF4 blocks requires WaveSpeed's xelerate runtime; the low-rank correction term is not part of the stock bitsandbytes NF4 format.

Related

License

Apache-2.0, inherited from Wan2.1.