Team Ai
Modelpublic

QuantFunc/Krea-2-QuantFunc-4bit

sourceHugging Faceotherupdated 6d agoView on Hugging Face
2likes5.1kdownloads
Model Card

<div align="center" style="margin-top: 50px;"> <a href="https://github.com/QuantFunc/ComfyUI-QuantFunc"> <img src="https://raw.githubusercontent.com/QuantFunc/ComfyUI-QuantFunc/main/assets/logo.webp" width="300" alt="QuantFunc logo"> </a> </div>

<p align="center"> ๐ŸŒ <a href="https://www.quantfunc.com/">Website</a> &nbsp;|&nbsp; ๐Ÿ™ <a href="https://github.com/QuantFunc/ComfyUI-QuantFunc">GitHub</a> &nbsp;|&nbsp; ๐Ÿค— <a href="https://huggingface.co/QuantFunc">Hugging Face</a> &nbsp;|&nbsp; ๐Ÿค– <a href="https://www.modelscope.cn/profile/QuantFunc">ModelScope</a> &nbsp;|&nbsp; ๐ŸŽจ <a href="https://civitai.com/user/quantfunc">Civitai</a> &nbsp;|&nbsp; ๐ŸŽฎ <a href="https://discord.gg/jCp9TpFWcn">Discord</a> </p>

Krea-2-QuantFunc-4bit

4x compression, quality held.

QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.

Showcase

All images below were generated by Krea-2-QuantFunc-4bit.

Portrait photographySci-fi scene
PortraitSci-fi
Impasto artWatercolor illustration
ImpastoWatercolor

Fast denoising, fast end-to-end too

High-VRAM

On an RTX 4090, same workflow and generation settings:

StageQuantFunc INT4FP8Speedup
Denoising1.6s5.0s3.13x
End-to-end2.5s6.5s2.6x
End-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.

Low-VRAM

On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup โ€” up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.

Swap one loader, keep your workflow

  1. 1.Install or update ComfyUI-QuantFunc.
  2. 2.Download the r128 or r32 weights.
  3. 3.Swap your model loader for the QuantFunc loader and pick the matching weight file.

Every other node, connection and generation parameter stays as-is.

Krea-2-Turbo is a distilled turbo model โ€” use fewer sampling steps and turn CFG off (guidance 1.0).

RTX 20-series through GB300, one build covers it all

Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

Choose a model

VariantFileSizeUse case
r128krea2-turbo-quantfunc-int4-r128.safetensors~8.3 GBQuality-first, recommended
r32krea2-turbo-quantfunc-int4-r32.safetensors~7.8 GBSize/VRAM-first

Loading

Weights use QuantFunc's own safetensors format โ€” load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint. (This repo declares library_name: diffusers purely so Hugging Face counts real downloads of the flat .safetensors files above โ€” see HF's download-stats docs โ€” not because the files load via the diffusers package.)

Source & license

This repository distributes derived quantized weights produced from:

Weights follow the Krea 2 Community License โ€” please read the original model's license terms before use.

Community