Team Ai
Modelpublic

akshan-main/tiny-qwenimage21-modular-pipe

sourceHugging Faceupdated 4d agoView on Hugging Face
0likes111downloads
Model Card

This is a modular diffusion pipeline built with ๐Ÿงจ Diffusers' modular pipeline framework.

Pipeline Type: QwenImage21AutoBlocks

Description: Auto Modular pipeline for text-to-image and image-conditioned generation using Qwen-Image 2.1.

  • โ€”for text-to-image generation, all you need to provide is prompt
  • โ€”for image-conditioned generation, you need to provide prompt and image (one image or a list)

This pipeline uses a 4-block architecture that can be customized and extended.

Example Usage

[TODO]

Pipeline Architecture

This modular pipeline is composed of the following blocks:

  1. 1.text_encoder (QwenImage21AutoTextEncoderStep)
  2. 2.Text encoder step that encodes the prompt, together with the condition images when there are any.
  3. 3.vae_encoder (QwenImage21AutoVaeEncoderStep)
  4. 4.VAE encoder step that encodes the condition images into their latent representations.
  5. 5.denoise (QwenImage21AutoCoreDenoiseStep)
  6. 6.Auto core denoise step that performs the denoising process.
  7. 7.decode (QwenImage21DecodeStep)
  8. 8.Step that decodes the latents to RGBA images and postprocesses them

Model Components

  1. 1.image_processor (VaeImageProcessor)
  2. 2.text_encoder (Qwen3VLForConditionalGeneration)
  3. 3.processor (Qwen3VLProcessor)
  4. 4.guider (ClassifierFreeGuidance)
  5. 5.vae (AutoencoderKLQwenImage21)
  6. 6.scheduler (FlowMatchEulerDiscreteScheduler)
  7. 7.transformer (QwenImage21Transformer2DModel)

Configuration Parameters

sample_sigmas (default: None): Default sampling grid of the checkpoint, used when sigmas is not passed.

Workflow Input Specification

<details> <summary><strong>text2image</strong></summary>

  • โ€”prompt (str): The prompt or prompts to guide image generation.

</details>

<details> <summary><strong>image_conditioned</strong></summary>

  • โ€”image (Image | list): Reference image(s) for denoising. Can be a single image or list of images.
  • โ€”prompt (str): The prompt or prompts to guide image generation.

</details>

Input/Output Specification

Inputs:

  • โ€”image (Image | list, optional): Reference image(s) for denoising. Can be a single image or list of images.
  • โ€”output_resolution (int, optional, defaults to 1024): Target side length used to derive the output size and to resize condition images.
  • โ€”prompt (str): The prompt or prompts to guide image generation.
  • โ€”negative_prompt (str, optional): The prompt or prompts not to guide the image generation.
  • โ€”generator (Generator, optional): Torch generator for deterministic generation.
  • โ€”num_images_per_prompt (int, optional, defaults to 1): The number of images to generate per prompt.
  • โ€”height (int, optional): The height in pixels of the generated image.
  • โ€”width (int, optional): The width in pixels of the generated image.
  • โ€”image_latents (list, optional): Normalized latents of each condition image. Can be generated from vae_encoder step.
  • โ€”latents (Tensor): Pre-generated noisy latents for image generation.
  • โ€”num_inference_steps (int): The number of denoising steps.
  • โ€”sigmas (list, optional): Custom sigmas for the denoising process.
  • โ€”use_kv_cache (bool, optional, defaults to True): Cache the text and condition-image keys and values after the first step. Valid because causal_condition modulates those tokens from t = 0, making their activations step-independent. Toggling it does not reproduce the same image bit-for-bit in reduced precision.
  • โ€”attention_kwargs (dict, optional): Additional kwargs for attention processors.
  • โ€”**denoiser_input_fields (None, optional): conditional model inputs for the denoiser: e.g. promptembeds, negativeprompt_embeds, etc.
  • โ€”output_type (str, optional, defaults to pil): Output format: 'pil', 'np', 'pt'.

Outputs:

  • โ€”images (list): Generated images.