Team Ai
Modelpublic

darrellbest/Qwen-Image-2.1-Text-Encoder-FP8

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
0likes32downloads
Model Card

Qwen-Image-2.1 Text Encoder: FP8

FP8 build of the text encoder shipped in Qwen/Qwen-Image-2.1 (text_encoder/), 9.9 GB instead of 16.3 GB. That encoder is stock Qwen/Qwen3-VL-8B-Instruct: all 750 tensors match, and every sampled tensor is byte-identical.

Format

compressed-tensors float-quantized, scheme FP8_DYNAMIC, made with llm-compressor 0.13.0 (no calibration needed):

PartPrecision
Every Linear layer of the language model (attention q/k/v/o, MLP gate/up/down)FP8 E4M3 weights (.weight) + one bf16 scale per output channel (.weight_scale, shape [out, 1]); activations dynamic per-token FP8 when run as W8A8
Embeddings, lm_head, norms, the whole vision towerbf16, unchanged

For a weight-only loader, the dequantized weight is simply weight.to(bf16) * weight_scale.

How it was tested

For a diffusion text encoder, what matters is what the image model reads: the last decoder layer's hidden states before the final norm, from the pipeline's own prompt templates. This was measured through diffusers' QwenImage21Pipeline.encode_prompt on held-out prompts (12 text-to-image, 8 image-edit on 2 held-out images), token by token against the bf16 original:

How it runsMean token cosine vs bf16Relative L2 error
Weight-only (FP8 weights dequantized, bf16 compute)0.97460.168
W8A8 (transformers + compressed-tensors, FP8 activations)0.95720.195

Same-seed 1024px images kept the layout of the bf16 ones, and rendered text stayed correct. For reference, GGUF Q80 measures 0.9938 / 0.070 on the same test: E4M3's 3-bit mantissa is coarser than 8-bit integers with a scale per 32 weights. If you need the closest match to bf16, use the Q80 in darrellbest/Qwen-Image-2.1-Text-Encoder-GGUF.

Use

python
from transformers import Qwen3VLForConditionalGeneration
te = Qwen3VLForConditionalGeneration.from_pretrained("darrellbest/Qwen-Image-2.1-Text-Encoder-FP8")

transformers decompresses the weights to bf16 on load (needs compressed-tensors), so this saves disk and download, not VRAM. A loader that keeps the FP8 weights (see Format) is what saves VRAM.

License: Apache-2.0, the license of Qwen3-VL-8B-Instruct. (The rest of Qwen-Image-2.1 is under the Qwen Research License; this repository contains only the encoder.)