Team Ai
Modelpublic

geceff/Wan2.2-Custom-Models-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
10likes3.2kdownloads
Model Card

Wan2.2 Custom GGUF & Specialized Models (Tesla T4 Optimized)

This repository provides highly optimized Wan2.2 Image-to-Video (I2V) GGUF and specialized custom models. These variants are tailored for running efficiently on memory-constrained environments, such as Google Colab equipped with an NVIDIA Tesla T4 GPU, while offering professional-grade motion extensions.


โš ๏ธ CRITICAL NOTICE: UN-UPDATED BASE MODEL WARNING

  • โ€”๐Ÿšจ Full-Size Base Models: Please note that the full-size raw models and non-quantized base files have NOW been updated in this repository.
  • โ€”๐Ÿ’ก Current Availability: Only the custom-compiled quants (GGUF), specialized LoRAs, Text Encoders (umt5_xxl_fp16 - umt5_xxl_fp8_e4m3fn_scaled), VAEs (Wan2_1_VAE_fp32 / wan_2.1_bf16), and specific FP8 integrated models are fully active and optimized for deployment. If you require raw unquantized BF16 weights, please wait for future repository syncs or utilize the available GGUF variants.

โšก Optimal Settings for ComfyUI

To achieve perfect video motion without artifacts or image degradation (preventing fried, burnt, or oversaturated visuals), we strongly recommend using the following parameters:

ParameterRecommended ValueNote
Total Sampling Steps4 - 12Absolute maximum ceiling is 12 total steps for Lightning / Distilled V2
CFG Scale1.0 - 2Crucial for preventing burnt images
High Noise Steps2, 4, 6, or 8To lock in strong motion. Can be split evenly (e.g., 8 steps total = 4 High / 4 Low) or 6 step total = 3 High / 3 low
Low Noise StepsDynamic (End Step: 4 - 12)CRITICAL: The target End Step for Low Noise must NEVER exceed the Total Sampling Steps!
Sampler / Schedulereuler + simpleStandard diffusion setup (Optionally, uni_pc can also be used for alternative fast-stepping)

๐Ÿ‘‘ Note for Higher Quality (Hybrid Workflow & Hardware Restrictions):

If you want to achieve higher visual fidelity and enhance micro-details, adopting a hybrid multi-pass approach is highly recommended. This strategy significantly sharpens fine details, effectively eliminates motion blur, and prevents fried visuals.

However, due to severe hardware VRAM limitations and Web GUI overhead, you MUST strictly adhere to the following setup configurations based on your execution environment:

๐Ÿ’ป 1. Via ComfyUI GUI (Web Interface Setup)
  • โ€”๐ŸŸข NVIDIA L4 (24GB VRAM) or higher: You can comfortably run high-tier configurations via the Web GUI with these setup options:
  • โ€”Standard High-Quality Setup: Use `Q8_H` (High Noise) + `Q8_H` (Low Noise) GGUF files.
  • โ€”Maximum Fidelity Option: Use `Q8_H` (High Noise) and chain it with `wan2.2_i2v_low_noise_14B_fp8_scaled.safetensors` as the final step to achieve ultimate sharpness and micro-details.
  • โ€”โš ๏ธ NVIDIA Tesla T4 (15GB VRAM - Free Tier GUI Limits): DO NOT use any models inside the `/diffusion_models` folder, nor any external `FP8` models placed outside in the root directory! Because the Web GUI consumes a massive amount of VRAM just to render its interface, available memory is extremely critical. Forcing these models via GUI will trigger an immediate OOM (Out of Memory) crash.
  • โ€”๐Ÿ“ Resolution Restrictions: Video dimensions must NOT exceed 480P in height (Max 480px) and NOT exceed 720P in width (Max 720px).
  • โ€”๐Ÿ›‘ Frame Limit: To remain stable, your generation length MUST NOT exceed 81-120(T4) frames max.
  • โ€”๐Ÿ›ก๏ธ Safe GUI Quantization Range: Your setup configurations via Web GUI are strictly restricted within the `Q4_K_M` to `Q8_H` range for both High Noise and Low Noise GGUF models.
๐Ÿš€ 2. Via Backdoor (Direct Code / Colab Forms Setup)
  • โ€”๐Ÿ”ฅ NVIDIA Tesla T4 (15GB VRAM - Unlocking Full Potential): By executing via the backend script directly, you bypass the heavy Web GUI memory overhead entirely, allowing you to forcefully squeeze maximum performance out of your T4 GPU!
  • โ€”The T4 Backdoor Formulas:
  • โ€”Ultimate Quality Setup: You can successfully execute the top-tier hybrid workflow: `Q4K_M` (High Noise) + `Q4K_M` (Low Noise).
  • โ€”Pro Option for Speed: If you want faster generation times with a minor trade-off, switch to `Q4_K` (High Noise) + `Q6_K` (Low Noise) or `Q4K_M.gguf` (High Noise) + `Q4K_M` (Low Noise). This delivers optimized speed while maintaining excellent visual quality compared to full high-quants.

๐Ÿ’พ Available Model Variants & Architecture.Reccomend onL4S or L4(slowly and short length totalfeame) or higher

Choose the right variant based on your creative workflow and VRAM configuration. All files are organized into dedicated subdirectories for pipeline flexibility:

๐ŸŽญ 1. Reccomend on L4S or higher Specialized Integrated FP8 Models (/diffusion_models)

These models feature pre-baked pipelines integrated with SVI (Stable Video Infinity) for continuous video synthesis and Consistent Face weights to prevent character distortion across frames.

  • โ€”`Wan2_2-I2V-A14B-HIGH_SVI_consistent_face_nsfw_fp8.safetensors`: Structural expert optimized for initial motion pathways, camera dynamics, and uncensored/free-form pipeline generations.
  • โ€”`Wan2_2-I2V-A14B-LOW_SVI_consistent_face_nsfw_fp8.safetensors`: Fine-tuning expert optimized for character preservation, facial structural lock, and detailed refinement.

โšก 2. Quantized Diffusion Models (Root Directory GGUF)

  • โ€”High Noise Quantizations (`wan2.2_i2v_high_noise_14B_...`): Best for creative, high-motion generation, and diverse camera movements. Available in: Q4_K_M, Q6_K_L, Q6_K, Q8_H, and fp8_scaled.
  • โ€”Low Noise Quantizations (`wan2.2_i2v_low_noise_14B_...`): Best for high fidelity, generation stability, and strictly adhering to the prompt or structural layout of your starting frame. Available in: Q4_K_M, Q6_K_L, Q6_K, Q8_H, and fp8_scaled.

๐Ÿงฉ 3. Modular Components

  • โ€”`/loras`: Contains raw targeted weights (high_noise and low_noise rank64 lightx2v 4-step) for modular multi-pass setups.
  • โ€”`/text_encoders`: Contains `umt5_xxl_fp16.safetensors` (11.4 GB) to maximize text prompt processing accuracy.
  • โ€”`/vae`: Contains `Wan2_1_VAE_fp32.safetensors` (508 MB) to prevent color degradation and artifacting during final video decoding.
  • โ€”Note: A lighter alternative `wan_2.1_bf16_vae.safetensors` is also placed in the root directory for extra VRAM safety during low-tier runs.