Team Ai
Modelpublic

diffusers-modular/krea2-edit

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
README.md58 linesDownload Raw Back to root
1---2library_name: diffusers3tags:4  - modular-diffusers5  - krea6  - image-to-image7  - image-editing8base_model: krea/Krea-2-Turbo9---10 11# Krea 2 reference-image edit — Modular Diffusers blocks12 13Custom [Modular Diffusers](https://huggingface.co/docs/diffusers/main/en/modular_diffusers/overview)14blocks that reproduce the [`ostris/Krea2OstrisEdit`](https://huggingface.co/ostris/Krea2OstrisEdit)15reference-image ("edit") workflow for **Krea 2**, loadable as remote code on top of stock `diffusers`.16 17```python18import torch19from transformers import Qwen3VLProcessor20from diffusers import ClassifierFreeGuidance21from diffusers.modular_pipelines import ModularPipelineBlocks22 23blocks = ModularPipelineBlocks.from_pretrained("diffusers-modular/krea2-edit", trust_remote_code=True)24pipe = blocks.init_pipeline("krea/Krea-2-Turbo")        # weights from the base repo25pipe.load_components(torch_dtype=torch.bfloat16)26pipe.update_components(processor=Qwen3VLProcessor.from_pretrained("Qwen/Qwen3-VL-4B-Instruct"))27pipe.update_components(guider=ClassifierFreeGuidance(guidance_scale=0.0, use_original_formulation=True))28pipe.to("cuda")29 30from PIL import Image31image = pipe(32    prompt="a white yeti with horns reading a book",33    image=Image.open("reference.png"),   # one or more reference images34    num_inference_steps=8, mu=1.15, output="images",35)[0]36```37 38## What it does39 40Reference images condition generation two ways (matching how the Ostris AI-Toolkit edit LoRAs train):41 421. **Qwen3-VL prompt embedding** — a coarse view of each reference is embedded into the text43   conditioning through the vision tower.442. **Clean VAE latents at flow time t=0** — each reference is VAE-encoded and appended to the45   transformer sequence as clean tokens on its own rotary frame axis, so the noisy image tokens46   attend to it at every block.47 48The bundled `Krea2Transformer2DModel` adds a small, backward-compatible `ref_seq_len` argument to the49Krea 2 transformer forward (t=0 modulation of the reference span; those tokens are excluded from the50predicted velocity). With `ref_seq_len=0` it is numerically identical to plain Krea 2 text-to-image.51 52## Files53 54- `block.py` — entry point (`Krea2EditBlocks`), referenced by `config.json`'s `auto_map`.55- `transformer_krea2.py` — the Krea 2 transformer (with the `ref_seq_len` edit path).56- `modular_blocks_krea2*.py`, `encoders.py`, `before_denoise.py`, `denoise.py`, `decoders.py`,57  `inputs.py`, `modular_pipeline.py` — the modular blocks.58