wavespeed/FLUX.1-dev-e4m3
06
1---2base_model: black-forest-labs/FLUX.1-dev3library_name: diffusers4license: other5license_name: flux-1-dev-non-commercial-license6license_link: https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md7pipeline_tag: text-to-image8tags:9- flux10- text-to-image11- quantized12- fp813- e4m314- diffusers15base_model_relation: quantized16---17# FLUX.1-dev-e4m318 19FP8 (e4m3) dynamically-quantized [FLUX.1-dev](https://huggingface.co/black-forest-labs/FLUX.1-dev),20saved as a complete `FluxPipeline`.21 22## What was changed23 24Every double and single transformer block of the `FluxTransformer2DModel` is25quantized to `e4m3_e4m3_dynamic` — `float8_e4m3fn` weights with dynamically26scaled `float8_e4m3fn` activations. The rest of the pipeline is unchanged: the27transformer's non-block tensors, the CLIP text encoder and the VAE stay in28fp16, and the T5 text encoder stays in bf16. The transformer shrinks from29~23.8 GB to ~12.0 GB.30 31This is the same recipe as32[`wavespeed/FLUX.1-dev-int8`](https://huggingface.co/wavespeed/FLUX.1-dev-int8)33with an fp8 rather than int8 numeric format. FP8 matmul needs Hopper (H100/H200)34or newer; on Ada and older the weights dequantize instead and you lose the speedup.35 36Quantization was done with WaveSpeed's `xelerate.ao.quantize`. Weights are37stored as pickled `.bin` shards, so loading requires `use_safetensors=False`.38 39## Usage40 41```python42import torch43from diffusers import FluxPipeline44 45pipe = FluxPipeline.from_pretrained(46 "wavespeed/FLUX.1-dev-e4m3",47 torch_dtype=torch.float16,48 use_safetensors=False,49).to("cuda")50```51 52## License53 54Derived from FLUX.1-dev, so the55[FLUX.1 \[dev\] Non-Commercial License](https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/main/LICENSE.md)56applies to these weights and to anything generated with them. Not for57commercial use.58 