lightx2v/Autoencoders
๐จ LightVAE
โก Efficient Video Autoencoder (VAE) Model Collection
From Official Models to Lightx2v Distilled Optimized Versions - Balancing Quality, Speed and Memory 
  
For VAE, the LightX2V team has conducted a series of deep optimizations, deriving two major series: LightVAE and LightTAE, which significantly reduce memory consumption and improve inference speed while maintaining high quality.
๐ก Core Advantages
<table> <tr> <td width="50%">
๐ Official VAE
Features: Highest Quality โญโญโญโญโญ
โ Best reconstruction accuracy โ Complete detail preservation โ Large memory usage (~8-12 GB) โ Slow inference speed
</td> <td width="50%">
๐ Open Source TAE Series
Features: Fastest Speed โกโกโกโกโก
โ Minimal memory usage (~0.4 GB) โ Extremely fast inference โ Average quality โญโญโญ โ Potential detail loss
</td> </tr> <tr> <td width="50%">
๐ฏ LightVAE Series (Our Optimization)
Features: Best Balanced Solution โ๏ธ
โ Uses Causal 3D Conv (same as official) โ Quality close to official โญโญโญโญ โ Memory reduced by ~50% (~4-5 GB) โ Speed increased by 2-3x โ Balances quality, speed, and memory ๐
</td> <td width="50%">
โก LightTAE Series (Our Optimization)
Features: Fast Speed + Good Quality ๐
โ Minimal memory usage (~0.4 GB) โ Extremely fast inference โ Quality close to official โญโญโญโญ โ Significantly surpasses open source TAE
</td> </tr> </table>
๐ฆ Available Models
๐ฏ Wan2.1 Series VAE
๐ฏ Wan2.2 Series VAE
๐ Wan2.1 Series Performance Comparison
- Precision: BF16
- Test Hardware: NVIDIA H100
Video Reconstruction (5s 81-frame video)
Video Generation
Task: s2v(speech to video) Model: seko-talk
<table> <tr> <td width="25%" align="center"> <strong>Wan2.1VAE</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/6l-P-3Hr9JKL3xgUyJXWJ.mp4"></video> </td> <td width="25%" align="center"> <strong>taew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/rcVHrCKB4nRAs2VSjJd2d.mp4"></video> </td> <td width="25%" align="center"> <strong>lighttaew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/Wq9p9Z7NDYwaKw4SqVbYT.mp4"></video> </td> <td width="25%" align="center"> <strong>lightvaew21</strong><br> <video controls autoplay muted width="100%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/NpKOzFcvsHzSFfFACzUKP.mp4"></video> </td> </tr> </table>
๐ Wan2.2 Series Performance Comparison
- Precision: BF16
- Test Hardware: NVIDIA H100
Video Reconstruction
Video Generation
Task: t2v(text to video) Model: Wan2.2-TI2V-5B
<table> <tr> <td width="33%" align="center"> <strong>Wan2.2VAE</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/KUY7Ifz9gFJqDjWga6A53.mp4"></video> </td> <td width="33%" align="center"> <strong>taew22</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/OYA8VfNlCvhBkjnOMl.mp4"></video> </td> <td width="33%" align="center"> <strong>lighttaew22</strong><br> <video controls autoplay width="95%" src="https://cdn-uploads.huggingface.co/production/uploads/680de13385293771bc57400b/gaHRr6uuAF0NlH4YlMbHO.mp4"></video> </td> </tr> </table>
๐ฏ Model Selection Recommendations
Selection by Use Case
<table> <tr> <td width="33%">
๐ Pursuing Best Quality
Recommended: Wan2.1_VAE / Wan2.2_VAE
- โ Official model, quality ceiling
- โ Highest reconstruction accuracy
- โ Suitable for final product output
- โ ๏ธ Large memory usage (~8-12 GB)
- โ ๏ธ Slow inference speed
</td> <td width="33%">
โ๏ธ Best Balance ๐
Recommended: `lightvaew2_1`
- โ Uses Causal 3D Conv (same as official)
- โ Excellent quality, close to official
- โ Memory reduced by ~50% (~4-5 GB)
- โ Speed increased by 2-3x
- โ Close to official quality โญโญโญโญ
Use Cases: Daily production, strongly recommended โญ
</td> <td width="33%">
โก Speed + Quality Balance โจ
Recommended: `lighttaew2_1` / `lighttaew2_2`
- โ Extremely low memory usage (~0.4 GB)
- โ Extremely fast inference
- โ Quality significantly surpasses open source TAE
- โ Close to official quality โญโญโญโญ
Use Cases: Development testing, rapid iteration
</td> </tr> </table>
๐ฅ Our Optimization Results Comparison
๐ Todo List
- [x] LightX2V integration
- [x] ComfyUI integration
- [ ] Training & Distillation Code
๐ Usage
Download VAE Models
# Download Wan2.1 official VAE
huggingface-cli download lightx2v/Autoencoders \
--local-dir ./models/vae/๐งช Video Reconstruction Test
We provide a standalone script vid_recon.py to test VAE models independently. This script reads a video, encodes it through VAE, then decodes it back to verify the reconstruction quality.
Script Location: LightX2V/lightx2v/models/video_encoders/hf/vid_recon.py
git clone https://github.com/ModelTC/LightX2V.git
cd LightX2V1. Test Official VAE (Wan2.1)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/Wan2.1_VAE.pth \
--model_type vaew2_1 \
--device cuda \
--dtype bfloat162. Test Official VAE (Wan2.2)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/Wan2.2_VAE.pth \
--model_type vaew2_2 \
--device cuda \
--dtype bfloat163. Test LightTAE (Wan2.1)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/lighttaew2_1.pth \
--model_type taew2_1 \
--device cuda \
--dtype bfloat164. Test LightTAE (Wan2.2)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/lighttaew2_2.pth \
--model_type taew2_2 \
--device cuda \
--dtype bfloat165. Test LightVAE (Wan2.1)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/lightvaew2_1.pth \
--model_type vaew2_1 \
--device cuda \
--dtype bfloat16 \
--use_lightvae6. Test TAE (Wan2.1)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/taew2_1.pth \
--model_type taew2_1 \
--device cuda \
--dtype bfloat167. Test TAE (Wan2.2)
python -m lightx2v.models.video_encoders.hf.vid_recon \
input_video.mp4 \
--checkpoint ./models/vae/taew2_2.pth \
--model_type taew2_1 \
--device cuda \
--dtype bfloat16Use in LightX2V
Specify the VAE path in the configuration file:
Using Official VAE Series:
{
"vae_path": "./models/vae/Wan2.1_VAE.pth"
}Using LightVAE Series:
{
"use_lightvae": true,
"vae_path": "./models/vae/lightvaew2_1.pth"
}Using LightTAE Series:
{
"use_tae": true,
"need_scaled": true,
"tae_path": "./models/vae/lighttaew2_1.pth"
}Using TAE Series:
{
"use_tae": true,
"tae_path": "./models/vae/taew2_1.pth"
}Then run the inference script:
cd LightX2V/scripts
bash wan/run_wan_i2v.sh # or other inference scriptsUse in ComfyUI
please refer to https://github.com/ModelTC/ComfyUI-LightVAE
โ ๏ธ Important Notes
1. Compatibility
- Wan2.1 series VAE only works with Wan2.1 backbone models
- Wan2.2 series VAE only works with Wan2.2 backbone models
- Do not mix different versions of VAE and backbone models
๐ Related Resources
Documentation Links
- LightX2V Quick Start: Quick Start Documentation
- Model Structure Description: Model Structure Documentation
- taeHV Project: GitHub - madebyollin/taeHV
Related Models
- Wan2.1 Backbone Models: Wan-AI Model Collection
- Wan2.2 Backbone Models: Wan-AI/Wan2.2-TI2V-5B
- LightX2V Optimized Models: lightx2v Model Collection
๐ค Community & Support
- GitHub Issues: https://github.com/ModelTC/LightX2V/issues
- HuggingFace: https://huggingface.co/lightx2v
- LightX2V Homepage: https://github.com/ModelTC/LightX2V
If you find this project helpful, please give us a โญ on GitHub
