darrellbest/Qwen-Image-2.1-Text-Encoder-FP8
Qwen-Image-2.1 Text Encoder: FP8
FP8 build of the text encoder shipped in Qwen/Qwen-Image-2.1 (text_encoder/), 9.9 GB instead of 16.3 GB. That encoder is stock Qwen/Qwen3-VL-8B-Instruct: all 750 tensors match, and every sampled tensor is byte-identical.
Format
compressed-tensors float-quantized, scheme FP8_DYNAMIC, made with llm-compressor 0.13.0 (no calibration needed):
For a weight-only loader, the dequantized weight is simply weight.to(bf16) * weight_scale.
How it was tested
For a diffusion text encoder, what matters is what the image model reads: the last decoder layer's hidden states before the final norm, from the pipeline's own prompt templates. This was measured through diffusers' QwenImage21Pipeline.encode_prompt on held-out prompts (12 text-to-image, 8 image-edit on 2 held-out images), token by token against the bf16 original:
Same-seed 1024px images kept the layout of the bf16 ones, and rendered text stayed correct. For reference, GGUF Q80 measures 0.9938 / 0.070 on the same test: E4M3's 3-bit mantissa is coarser than 8-bit integers with a scale per 32 weights. If you need the closest match to bf16, use the Q80 in darrellbest/Qwen-Image-2.1-Text-Encoder-GGUF.
Use
from transformers import Qwen3VLForConditionalGeneration
te = Qwen3VLForConditionalGeneration.from_pretrained("darrellbest/Qwen-Image-2.1-Text-Encoder-FP8")transformers decompresses the weights to bf16 on load (needs compressed-tensors), so this saves disk and download, not VRAM. A loader that keeps the FP8 weights (see Format) is what saves VRAM.
License: Apache-2.0, the license of Qwen3-VL-8B-Instruct. (The rest of Qwen-Image-2.1 is under the Qwen Research License; this repository contains only the encoder.)
