Team Ai
Modelpublic

lightx2v/Autoencoders

sourceHugging Faceapache-2.0updated 11mo agoView on Hugging Face
132likes7.2kdownloads
Model Card

๐ŸŽจ LightVAE

โšก Efficient Video Autoencoder (VAE) Model Collection

From Official Models to Lightx2v Distilled Optimized Versions - Balancing Quality, Speed and Memory img_lightx2v


![๐Ÿค— HuggingFace](https://huggingface.co/lightx2v) ![GitHub](https://github.com/ModelTC/LightX2V) ![License](LICENSE)


For VAE, the LightX2V team has conducted a series of deep optimizations, deriving two major series: LightVAE and LightTAE, which significantly reduce memory consumption and improve inference speed while maintaining high quality.

๐Ÿ’ก Core Advantages

<table> <tr> <td width="50%">

๐Ÿ“Š Official VAE

Features: Highest Quality โญโญโญโญโญ

โœ… Best reconstruction accuracy โœ… Complete detail preservation โŒ Large memory usage (~8-12 GB) โŒ Slow inference speed

</td> <td width="50%">

๐Ÿš€ Open Source TAE Series

Features: Fastest Speed โšกโšกโšกโšกโšก

โœ… Minimal memory usage (~0.4 GB) โœ… Extremely fast inference โŒ Average quality โญโญโญ โŒ Potential detail loss

</td> </tr> <tr> <td width="50%">

๐ŸŽฏ LightVAE Series (Our Optimization)

Features: Best Balanced Solution โš–๏ธ

โœ… Uses Causal 3D Conv (same as official) โœ… Quality close to official โญโญโญโญ โœ… Memory reduced by ~50% (~4-5 GB) โœ… Speed increased by 2-3x โœ… Balances quality, speed, and memory ๐Ÿ†

</td> <td width="50%">

โšก LightTAE Series (Our Optimization)

Features: Fast Speed + Good Quality ๐Ÿ†

โœ… Minimal memory usage (~0.4 GB) โœ… Extremely fast inference โœ… Quality close to official โญโญโญโญ โœ… Significantly surpasses open source TAE

</td> </tr> </table>


๐Ÿ“ฆ Available Models

๐ŸŽฏ Wan2.1 Series VAE

Model NameTypeArchitectureDescription
Wan2.1_VAEOfficial VAECausal Conv3DWan2.1 official video VAE model<br>Highest quality, large memory, slow speed
taew2_1Open Source Small AEConv2DOpen source model based on taeHV<br>Small memory, fast speed, average quality
`lighttaew2_1`LightTAE SeriesConv2DOur distilled optimized version based on `taew2_1`<br>Small memory, fast speed, quality close to official โœจ
`lightvaew2_1`LightVAE SeriesCausal Conv3DOur pruned 75% on WanVAE2.1 architecture then trained+distilled<br>Best balance: high quality + low memory + fast speed ๐Ÿ†

๐ŸŽฏ Wan2.2 Series VAE

Model NameTypeArchitectureDescription
Wan2.2_VAEOfficial VAECausal Conv3DWan2.2 official video VAE model<br>Highest quality, large memory, slow speed
taew2_2Open Source Small AEConv2DOpen source model based on taeHV<br>Small memory, fast speed, average quality
`lighttaew2_2`LightTAE SeriesConv2DOur distilled optimized version based on `taew2_2`<br>Small memory, fast speed, quality close to official โœจ

๐Ÿ“Š Wan2.1 Series Performance Comparison

  • โ€”Precision: BF16
  • โ€”Test Hardware: NVIDIA H100

Video Reconstruction (5s 81-frame video)

SpeedWan2.1_VAEtaew2_1lighttaew2_1lightvaew2_1
Encode Speed4.1721 s0.3956 s0.3956 s1.5014s
Decode Speed5.4649 s0.2463 s0.2463 s2.0697s
GPU MemoryWan2.1_VAEtaew2_1lighttaew2_1lightvaew2_1
Encode Memory8.4954 GB0.00858 GB0.00858 GB4.7631 GB
Decode Memory10.1287 GB0.41199 GB0.41199 GB5.5673 GB

Video Generation

Task: s2v(speech to video) Model: seko-talk

<table> <tr> <td width="25%" align="center"> <strong>Wan2.1VAE</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/6l-P-3Hr9JKL3xgUyJXWJ.mp4"></video> </td> <td width="25%" align="center"> <strong>taew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/rcVHrCKB4nRAs2VSjJd2d.mp4"></video> </td> <td width="25%" align="center"> <strong>lighttaew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/Wq9p9Z7NDYwaKw4SqVbYT.mp4"></video> </td> <td width="25%" align="center"> <strong>lightvaew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/NpKOzFcvsHzSFfFACzUKP.mp4"></video> </td> </tr> </table>

๐Ÿ“Š Wan2.2 Series Performance Comparison

  • โ€”Precision: BF16
  • โ€”Test Hardware: NVIDIA H100

Video Reconstruction

SpeedWan2.2_VAEtaew2_2lighttaew2_2
Encode Speed1.1369s0.3499 s0.3499 s
Decode Speed3.1268 s0.0891 s0.0891 s
GPU MemoryWan2.2_VAEtaew2_2lighttaew2_2
Encode Memory6.1991 GB0.0064 GB0.0064 GB
Decode Memory12.3487 GB0.4120 GB0.4120 GB

Video Generation

Task: t2v(text to video) Model: Wan2.2-TI2V-5B

<table> <tr> <td width="33%" align="center"> <strong>Wan2.2VAE</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/KUY7Ifz9gFJqDjWga6A53.mp4"></video> </td> <td width="33%" align="center"> <strong>taew22</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/OYA8VfNlCvhBkjnOMl.mp4"></video> </td> <td width="33%" align="center"> <strong>lighttaew22</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/gaHRr6uuAF0NlH4YlMbHO.mp4"></video> </td> </tr> </table>

๐ŸŽฏ Model Selection Recommendations

Selection by Use Case

<table> <tr> <td width="33%">

๐Ÿ† Pursuing Best Quality

Recommended: Wan2.1_VAE / Wan2.2_VAE

  • โ€”โœ… Official model, quality ceiling
  • โ€”โœ… Highest reconstruction accuracy
  • โ€”โœ… Suitable for final product output
  • โ€”โš ๏ธ Large memory usage (~8-12 GB)
  • โ€”โš ๏ธ Slow inference speed

</td> <td width="33%">

โš–๏ธ Best Balance ๐Ÿ†

Recommended: `lightvaew2_1`

  • โ€”โœ… Uses Causal 3D Conv (same as official)
  • โ€”โœ… Excellent quality, close to official
  • โ€”โœ… Memory reduced by ~50% (~4-5 GB)
  • โ€”โœ… Speed increased by 2-3x
  • โ€”โœ… Close to official quality โญโญโญโญ

Use Cases: Daily production, strongly recommended โญ

</td> <td width="33%">

โšก Speed + Quality Balance โœจ

Recommended: `lighttaew2_1` / `lighttaew2_2`

  • โ€”โœ… Extremely low memory usage (~0.4 GB)
  • โ€”โœ… Extremely fast inference
  • โ€”โœ… Quality significantly surpasses open source TAE
  • โ€”โœ… Close to official quality โญโญโญโญ

Use Cases: Development testing, rapid iteration

</td> </tr> </table>

๐Ÿ”ฅ Our Optimization Results Comparison

ComparisonOpen Source TAE**LightTAE (Ours)**Official VAE**LightVAE (Ours)**
ArchitectureConv2DConv2DCausal Conv3DCausal Conv3D
Memory UsageMinimal (~0.4 GB)Minimal (~0.4 GB)Large (~8-12 GB)Medium (~4-5 GB)
Inference SpeedExtremely Fast โšกโšกโšกโšกโšกExtremely Fast โšกโšกโšกโšกโšกSlow โšกโšกFast โšกโšกโšกโšก
Generation QualityAverage โญโญโญClose to Official โญโญโญโญHighest โญโญโญโญโญClose to Official โญโญโญโญ

๐Ÿ“‘ Todo List

  • โ€”[x] LightX2V integration
  • โ€”[x] ComfyUI integration
  • โ€”[ ] Training & Distillation Code

๐Ÿš€ Usage

Download VAE Models

bash
# Download Wan2.1 official VAE
huggingface-cli download lightx2v/Autoencoders \
    --local-dir ./models/vae/

๐Ÿงช Video Reconstruction Test

We provide a standalone script vid_recon.py to test VAE models independently. This script reads a video, encodes it through VAE, then decodes it back to verify the reconstruction quality.

Script Location: LightX2V/lightx2v/models/video_encoders/hf/vid_recon.py

bash
git clone https://github.com/ModelTC/LightX2V.git
cd LightX2V

1. Test Official VAE (Wan2.1)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/Wan2.1_VAE.pth \
    --model_type vaew2_1 \
    --device cuda \
    --dtype bfloat16

2. Test Official VAE (Wan2.2)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/Wan2.2_VAE.pth \
    --model_type vaew2_2 \
    --device cuda \
    --dtype bfloat16

3. Test LightTAE (Wan2.1)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/lighttaew2_1.pth \
    --model_type taew2_1 \
    --device cuda \
    --dtype bfloat16

4. Test LightTAE (Wan2.2)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/lighttaew2_2.pth \
    --model_type taew2_2 \
    --device cuda \
    --dtype bfloat16

5. Test LightVAE (Wan2.1)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/lightvaew2_1.pth \
    --model_type vaew2_1 \
    --device cuda \
    --dtype bfloat16 \
    --use_lightvae

6. Test TAE (Wan2.1)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/taew2_1.pth \
    --model_type taew2_1 \
    --device cuda \
    --dtype bfloat16

7. Test TAE (Wan2.2)

bash
python -m lightx2v.models.video_encoders.hf.vid_recon \
    input_video.mp4 \
    --checkpoint ./models/vae/taew2_2.pth \
    --model_type taew2_1 \
    --device cuda \
    --dtype bfloat16

Use in LightX2V

Specify the VAE path in the configuration file:

Using Official VAE Series:

json
{

    "vae_path": "./models/vae/Wan2.1_VAE.pth"
}

Using LightVAE Series:

json
{
    "use_lightvae": true,
    "vae_path": "./models/vae/lightvaew2_1.pth"
}

Using LightTAE Series:

json
{
    "use_tae": true,
    "need_scaled": true,
    "tae_path": "./models/vae/lighttaew2_1.pth"
}

Using TAE Series:

json
{
    "use_tae": true,
    "tae_path": "./models/vae/taew2_1.pth"
}

Then run the inference script:

bash
cd LightX2V/scripts
bash wan/run_wan_i2v.sh  # or other inference scripts

Use in ComfyUI

please refer to https://github.com/ModelTC/ComfyUI-LightVAE

โš ๏ธ Important Notes

1. Compatibility

  • โ€”Wan2.1 series VAE only works with Wan2.1 backbone models
  • โ€”Wan2.2 series VAE only works with Wan2.2 backbone models
  • โ€”Do not mix different versions of VAE and backbone models

๐Ÿ“š Related Resources

Documentation Links

Related Models


๐Ÿค Community & Support

  • โ€”GitHub Issues: https://github.com/ModelTC/LightX2V/issues
  • โ€”HuggingFace: https://huggingface.co/lightx2v
  • โ€”LightX2V Homepage: https://github.com/ModelTC/LightX2V

If you find this project helpful, please give us a โญ on GitHub