Team Ai
Modelpublic

alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union

sourceHugging Faceotherupdated 17d agoView on Hugging Face
55likes36kdownloads
Model Card

Qwen-Image-2.1-Fun-Controlnet-Union

![Github](https://github.com/aigc-apps/VideoX-Fun)

Overview

Qwen-Image-2.1-Fun-Controlnet-Union is a ControlNet-Union branch for [Qwen-Image 2.1](https://github.com/aigc-apps/VideoX-Fun) (the flow-matching text-to-image DiT). A single checkpoint drives 8 structural control conditions (Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose, Scribble) and image inpainting, without per-condition weights. The checkpoint holds only the control branch (control_img_in plus 16 control_blocks, about 7.0 GB) and is loaded on top of the base Qwen-Image 2.1 transformer.

Model Card

NameDescription
Qwen-Image-2.1-Fun-Controlnet-Union.safetensorsControlNet-Union branch weights for Qwen-Image 2.1. Contains only the control branch (control_img_in + 16 control_blocks, about 7.0 GB); loaded with strict=False on top of the base Qwen-Image 2.1 transformer. One checkpoint covers 8 control conditions and image inpainting.

Model Features

  • โ€”Union control over 8 conditions: one checkpoint handles Canny, Depth, Grayscale, HED, Lineart, MLSD, Pose and Scribble control images for text-to-image generation โ€” no per-condition checkpoint switching.
  • โ€”Dense control injection: the control branch attaches a skip to every 2nd of the 32 transformer blocks (control_layers = [0, 2, 4, โ€ฆ, 30], 16 injection points). Each control skip is added back to the main branch through zero-gated before_proj / after_proj projections, giving tight structural adherence while the base model stays frozen.
  • โ€”Control and inpainting share one branch: the control input is widened to control_in_dim = 129 โ€” control latents (64) | mask (1) | masked-image latents (64). For pure control the mask / masked-image channels are zero-padded; for inpainting the same branch re-draws the masked region from the prompt. The two can also be combined โ€” a control image and a mask are fed together, so the re-drawn region follows both the prompt and the given structure.
  • โ€”CFG-distilled fast sampling: the standalone example scripts run with guidance_scale = 1.0 (single forward pass per step, no classifier-free guidance needed).
  • โ€”control_context_scale scales every control skip before it is added to the main branch: 1.0 is the strongest control (used for all results below), lower values weaken the guidance, 0.0 switches the control branch off.
  • โ€”Prompt-friendly: write a prompt that describes the whole target image; the masked region is conveyed by the mask channel, not by the text. Detailed prompts give better stability.
  • โ€”Qwen-Image 2.1 encodes the prompt (and any condition image) with a Qwen3-VL text encoder + processor, and its VAE decodes to RGBA, so every preview is saved as PNG.

Supported control conditions

ConditionControl signal
CannyCanny edge map
DepthMonocular depth map
GrayscaleGrayscale (luminance) image
HEDHED edge detection map
LineartLine-art extraction
MLSDLine-segment detection map
PoseDWPose skeleton
ScribbleFree-hand / sketch lines

Any ordinary RGB control image at the target canvas works; the model tolerates different line thickness, thresholds and crops.

Results

All samples below are generated with num_inference_steps = 40, control_context_scale = 1.0, seed 43. In each column the top row is the control image, the bottom row is the output.

<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr><td>Canny</td><td>Depth</td><td>Grayscale</td><td>HED</td><td>Lineart</td><td>MLSD</td><td>Pose</td><td>Scribble</td></tr> <tr> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control200025069.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control500005389.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control100014291.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control600009388.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control400003435.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control400027496.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control400000931.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/control1300000494.png" width="100%"></td> </tr> <tr> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control200025069.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control500005389.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control100014291.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control600009388.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control400003435.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control400027496.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control400000931.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/control1300000494.png" width="100%"></td> </tr> </table>

Inpainting (+ control)

A masked region of the source image is re-drawn from the prompt while the rest of the frame is preserved. The mask image is white where the content should be re-generated and black where it should be kept. Because control and inpainting share the same branch, a control image (here a DWPose skeleton) is fed together with the mask, so the re-drawn region also follows the given pose.

<table border="0" style="width: 100%; text-align: left; margin-top: 20px;"> <tr><td>Source image</td><td>Mask</td><td>Pose control</td><td>Inpaint output</td></tr> <tr> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/inpaintsource.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/inpaintmask.png" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/asset/inpaint_control.jpg" width="100%"></td> <td><img src="https://huggingface.co/alibaba-pai/Qwen-Image-2.1-Fun-Controlnet-Union/resolve/main/results/inpaint.png" width="100%"></td> </tr> </table>

Inference

Go to the VideoX-Fun repository for more details.

Please clone the VideoX-Fun repository and create the required directories:

sh
# Clone the code
git clone https://github.com/aigc-apps/VideoX-Fun.git

# Enter VideoX-Fun's directory
cd VideoX-Fun

# Create model directories
mkdir -p models/Diffusion_Transformer

Then download the base Qwen-Image 2.1 model and this checkpoint into models/Diffusion_Transformer.

๐Ÿ“ฆ models/
โ”œโ”€โ”€  Diffusion_Transformer/
โ”‚   โ”œโ”€โ”€ ๐Ÿ“‚ Qwen-Image-2.1/
โ”‚   โ””โ”€โ”€ ๐Ÿ“‚ Qwen-Image-2.1-Fun-Controlnet-Union/
โ”‚       โ””โ”€โ”€  Qwen-Image-2.1-Fun-Controlnet-Union.safetensors

Then edit the settings at the top of examples/qwenimage21_fun/predict_t2i_control.py (or predict_i2i_inpaint.py for inpainting) and run it.

python
model_name          = "models/Diffusion_Transformer/Qwen-Image-2.1"
config_path         = "config/qwenimage21/qwenimage21_control.yaml"
transformer_path    = "models/Diffusion_Transformer/Qwen-Image-2.1-Fun-Controlnet-Union/Qwen-Image-2.1-Fun-Controlnet-Union.safetensors"
control_image       = "asset/pose.jpg"
# inpaint only:
inpaint_image       = "asset/8.png"
mask_image          = "asset/mask.png"
prompt              = "your prompt describing the whole target image"
sh
python examples/qwenimage21_fun/predict_t2i_control.py

Notes:

  • โ€”config_path must be config/qwenimage21/qwenimage21_control.yaml. It builds the control branch exactly as the checkpoint expects (control_layers: [0, 2, 4, โ€ฆ, 30], control_in_dim: 129); a mismatched config silently drops or misplaces control weights and produces wrong outputs.
  • โ€”For pure control (no inpaint input) the pipeline zero-pads the mask / masked-image channels, so this inpaint-capable checkpoint still runs plain Canny/Depth/โ€ฆ control correctly.
  • โ€”The control checkpoint carries only the control branch; the base Qwen-Image 2.1 weights must be present in model_name.
  • โ€”control_context_scale = 1.0 is the value the adapter expects; use lower values to loosen the structural constraint.
  • โ€”sample_size sets the output canvas (e.g. [1728, 992]); keep both sides as multiples of 16 so the control map is not distorted.
  • โ€”use_kv_cache = True caches the text / condition-image keys after the first denoising step for a speedup at fixed resolution.
  • โ€”Memory: the transformer plus the Qwen3-VL text encoder do not fit a single consumer GPU fully loaded; use model_group_offload (fastest) or model_cpu_offload_and_qfloat8 on a single high-memory GPU.

License

This model is a derivative of Qwen-Image 2.1 and is released under the Qwen Research License. Please read the license carefully before use.