BasinShapers/flux2-klein-4b-webgpu
FLUX.2 [klein] 4B for WebGPU
Quantized weights of [FLUX.2 [klein] 4B](https://huggingface.co/black-forest-labs/FLUX.2-klein-4B), packed for text-to-image generation entirely in the browser with the Kaminos WebGPU inference kit. The page downloads about 3.9 GB once (everything except the embedding table, from which it fetches only the rows a prompt uses), caches it, and generates on the visitor's own GPU. No server-side compute is involved.
[Try it in your browser](https://lyonsno.github.io/kaminos/inference-kit/klein/) · Source
In Chrome, a 512 × 512 image (4 steps) takes 6–7 seconds on an Apple M4 Max and about 20 seconds on a 16 GB Apple M2 Pro. The page also offers 768 and 1024, and frees each stage's working memory before the next, so a 16 GB Mac generates at 1024 × 1024 without swapping.
Contents
Each folder has a manifest.json listing every tensor's shape, format, byte offset, and the SHA-256 of each file. Quantization is weight-only. Activations and the residual stream stay in floating point.
Changes from the original
- Weights quantized and repacked into per-block files as described above.
- Text encoder truncated to the layers the FLUX.2 [klein] pipeline uses.
- Projections that share an input are fused (for example query, key and value).
- 3×3 VAE convolution weights reordered for channels-last execution.
License
Apache License 2.0, the same license as the original model. FLUX.2 [klein] 4B is by Black Forest Labs. Its text encoder is Qwen3-4B by the Qwen team, also Apache 2.0.
