Team Ai
Modelpublic

Efficient-Large-Model/SANA-Video_2B_720p_diffusers

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
6likes
README.md210 linesDownload Raw Back to root
1---2license: apache-2.03library_name: sana, sana-video4tags:5- text-to-video6- SANA-Video7- 720p_5s_pretrained_model8- BF169- diffusion10language:11- en12- zh13base_model:14- Efficient-Large-Model/SANA-Video_2B_720p_diffusers15pipeline_tag: text-to-video16---17<p align="center" style="border-radius: 10px">18  <img src="https://cdn-uploads.huggingface.co/production/uploads/645b5b09bc7518912e1f9733/N0VlE-y1pau-4O1RlijQd.png" width="98%" alt="logo"/>19</p>20 21<div style="display:flex;justify-content: center">22  <a href="https://hf.co/collections/Efficient-Large-Model/sana-video"><img src="https://img.shields.io/static/v1?label=Weights&message=Huggingface&color=yellow"></a> &ensp;23  <a href="https://github.com/NVlabs/Sana"><img src="https://img.shields.io/static/v1?label=Code&message=Github&color=blue&logo=github"></a> &ensp;24  <a href="https://nvlabs.github.io/Sana/Video/"><img src="https://img.shields.io/static/v1?label=Project&message=Github&color=blue&logo=github-pages"></a> &ensp;25  <a href="https://arxiv.org/pdf/2509.24695"><img src="https://img.shields.io/static/v1?label=Arxiv&message=SANA-Video&color=red&logo=arxiv"></a> &ensp;26</div>27 28 29# ๐Ÿฑ SANA-Video Model Card30 31<!-- <div align="center">32  <a href="https://www.youtube.com/watch?v=nI_Ohgf8eOU" target="_blank">33    <img src="https://img.youtube.com/vi/nI_Ohgf8eOU/0.jpg" alt="Demo Video of SANA-Video" style="width: 48%; display: block; margin: 0 auto; display: inline-block;">34  </a>35  <a href="https://www.youtube.com/watch?v=OOZzkirgsAc" target="_blank">36    <img src="https://img.youtube.com/vi/OOZzkirgsAc/0.jpg" alt="Demo Video of SANA-Video" style="width: 48%; display: block; margin: 0 auto; display: inline-block;">37  </a>38</div> -->39 40 41SANA-Video is a small, ultra-efficient diffusion model designed for rapid generation of high-quality, minute-long videos at resolutions up to 720ร—1280.42 43Key innovations and efficiency drivers include:44 45(1) **Linear DiT**: Leverages linear attention as the core operation, offering significantly more efficiency than vanilla attention when processing the massive number of tokens required for video generation.46 47(2) **Constant-Memory KV Cache for Block Linear Attention**: Implements a block-wise autoregressive approach that uses the cumulative properties of linear attention to maintain global context at a fixed memory cost, eliminating the traditional KV cache bottleneck and enabling efficient, minute-long video synthesis.48 49SANA-Video achieves exceptional efficiency and cost savings: its training cost is only **1%** of MovieGen's (**12 days on 64 H100 GPUs**). Compared to modern state-of-the-art small diffusion models (e.g., Wan 2.1 and SkyReel-V2), SANA-Video maintains competitive performance while being **16ร—** faster in measured latency.50SANA-Video is deployable on RTX 5090 GPUs, accelerating the inference speed for a 5-second 720p video from 71s down to 29s (2.4ร— speedup), setting a new standard for low-cost, high-quality video generation.51 52Source code is available at https://github.com/NVlabs/Sana.53 54# ๐Ÿฑ How to Inference55 56Refer to: https://github.com/NVlabs/Sana/blob/main/asset/docs/sana_video.md#1-inference-with-txt-file57 58# diffusers pipeline59 60```python61"""Sana Video + LTX2 Refiner: Stage 1 generate latent โ†’ Stage 2 refine (3 steps)."""62 63import gc64import torch65from diffusers import SanaVideoPipeline, FlowMatchEulerDiscreteScheduler66from diffusers.pipelines.ltx2 import LTX2Pipeline, LTX2LatentUpsamplePipeline67from diffusers.pipelines.ltx2.latent_upsampler import LTX2LatentUpsamplerModel68from diffusers.pipelines.ltx2.utils import STAGE_2_DISTILLED_SIGMA_VALUES69from diffusers.pipelines.ltx2.export_utils import encode_video70 71device = "cuda"72dtype = torch.bfloat1673prompt = "A cat walking on the grass, facing the camera."74negative_prompt = "A chaotic sequence with misshapen, deformed limbs in heavy motion blur, sudden disappearance, jump cuts, jerky movements, rapid shot changes, frames out of sync, inconsistent character shapes, temporal artifacts, jitter, and ghosting effects, creating a disorienting visual experience."75motion_score = 3076height, width, frames, frame_rate = 704, 1280, 81, 16.077seed = 4278 79# โ”€โ”€ Load all models โ”€โ”€80sana_pipe = SanaVideoPipeline.from_pretrained(81    "Efficient-Large-Model/SANA-Video_2B_720p_diffusers", torch_dtype=dtype,82)83sana_pipe.text_encoder.to(dtype)84sana_pipe.enable_model_cpu_offload()85 86ltx_pipe = LTX2Pipeline.from_pretrained("Lightricks/LTX-2", torch_dtype=dtype)87ltx_pipe.load_lora_weights(88    "Lightricks/LTX-2", adapter_name="stage_2_distilled",89    weight_name="ltx-2-19b-distilled-lora-384.safetensors",90)91ltx_pipe.set_adapters("stage_2_distilled", 1.0)92ltx_pipe.vae.enable_tiling()93ltx_pipe.enable_model_cpu_offload()94 95latent_upsampler = LTX2LatentUpsamplerModel.from_pretrained(96    "Lightricks/LTX-2", subfolder="latent_upsampler", torch_dtype=dtype,97)98upsample_pipe = LTX2LatentUpsamplePipeline(vae=ltx_pipe.vae, latent_upsampler=latent_upsampler)99upsample_pipe.enable_model_cpu_offload(device=device)100 101# โ”€โ”€ Stage 1: Sana Video โ”€โ”€102video_latent = sana_pipe(103    prompt=prompt + f" motion score: {motion_score}.", negative_prompt=negative_prompt,104    height=height, width=width, frames=frames,105    guidance_scale=6.0, num_inference_steps=50,106    generator=torch.Generator(device=device).manual_seed(seed),107    output_type="latent", return_dict=True,108).frames109 110del sana_pipe; gc.collect(); torch.cuda.empty_cache()111 112# โ”€โ”€ Stage 1.5: Latent Upsample (2x spatial) โ”€โ”€113video_latent = upsample_pipe(114    latents=video_latent.to(device=device, dtype=dtype),115    latents_normalized=True,116    height=height, width=width, num_frames=frames,117    output_type="latent", return_dict=False,118)[0]119latents_mean = ltx_pipe.vae.latents_mean.view(1, -1, 1, 1, 1).to(video_latent.device, video_latent.dtype)120latents_std = ltx_pipe.vae.latents_std.view(1, -1, 1, 1, 1).to(video_latent.device, video_latent.dtype)121video_latent = (video_latent - latents_mean) * ltx_pipe.vae.config.scaling_factor / latents_std122 123# โ”€โ”€ Stage 2: LTX2 Refine โ”€โ”€124packed = LTX2Pipeline._pack_latents(125    video_latent.to(device=device, dtype=dtype),126    patch_size=ltx_pipe.transformer_spatial_patch_size,127    patch_size_t=ltx_pipe.transformer_temporal_patch_size,128)129_, _, lF, lH, lW = video_latent.shape130pH, pW, pT = lH * ltx_pipe.vae_spatial_compression_ratio, lW * ltx_pipe.vae_spatial_compression_ratio, (lF - 1) * ltx_pipe.vae_temporal_compression_ratio + 1131 132dur = pT / frame_rate133audio_frames = round(dur * ltx_pipe.audio_sampling_rate / ltx_pipe.audio_hop_length / ltx_pipe.audio_vae_temporal_compression_ratio)134nch = ltx_pipe.audio_vae.config.latent_channels135mel = ltx_pipe.audio_vae.config.mel_bins // ltx_pipe.audio_vae_mel_compression_ratio136audio_latent = (137    ltx_pipe.audio_vae.latents_mean.unsqueeze(0).unsqueeze(0)138    .expand(1, audio_frames, nch * mel).to(dtype=dtype, device=device).contiguous()139    .unflatten(2, (nch, mel)).permute(0, 2, 1, 3).contiguous()140)141 142del video_latent; gc.collect(); torch.cuda.empty_cache()143 144video, _ = ltx_pipe(145    latents=packed, audio_latents=audio_latent,146    prompt=prompt, negative_prompt=negative_prompt,147    height=pH, width=pW, num_frames=pT,148    num_inference_steps=3,149    noise_scale=STAGE_2_DISTILLED_SIGMA_VALUES[0],150    sigmas=STAGE_2_DISTILLED_SIGMA_VALUES,151    guidance_scale=1.0, frame_rate=frame_rate,152    generator=torch.Generator(device=device).manual_seed(seed),153    output_type="np", return_dict=False,154)155 156video = torch.from_numpy((video * 255).round().astype("uint8"))157encode_video(video[0], fps=frame_rate, audio=None, audio_sample_rate=None, output_path="sana_ltx2_refined.mp4")158```159 160### Model Description161 162- **Developed by:** NVIDIA, Sana163- **Model type:** Efficient Video Generation with Block Linear Diffusion Transformer164- **Model size:** 2B parameters165- **Model precision:** torch.bfloat16 (BF16)166- **Model resolution:** This model is developed to generate 720p resolution 81(5s) frames videos with multi-scale heigh and width.167- **Model Description:** This is a model that can be used to generate and modify videos based on text prompts. 168It is a Linear Diffusion Transformer that uses LTX2-vae one 32x32x8 spatial-temporal-compressed latent feature encoder ([LTX2](https://github.com/Lightricks/LTX-2)).169- **Resources for more information:** Check out our [GitHub Repository](https://github.com/NVlabs/Sana) and the [SANA-Video report on arXiv](https://arxiv.org/pdf/2509.24695).170 171### Model Sources172 173For research purposes, we recommend our `generative-models` Github repository (https://github.com/NVlabs/Sana), which is more suitable for both training and inference174- **Repository:** https://github.com/NVlabs/Sana175- **Guidance:** https://github.com/NVlabs/Sana/asset/docs/sana_video.md176 177## License/Terms of Use178 179This model is released under the [Apache License 2.0](./LICENSE).180 181## Uses182 183### Direct Use184 185The model is intended for research purposes only. Possible research areas and tasks include186 187- Generation of artworks and use in design and other artistic processes.188- Applications in educational or creative tools.189- Research on generative models.190- Safe deployment of models which have the potential to generate harmful content.191 192- Probing and understanding the limitations and biases of generative models.193 194Excluded uses are described below.195 196### Out-of-Scope Use197 198The model was not trained to be factual or true representations of people or events, and therefore using the model to generate such content is out-of-scope for the abilities of this model.199 200## Limitations and Bias201 202### Limitations203 204- The model does not achieve perfect photorealism205- The model cannot render complex legible text206- fingers, .etc in general may not be generated properly.207- The autoencoding part of the model is lossy.208 209### Bias210While the capabilities of video generation models are impressive, they can also reinforce or exacerbate social biases.