Team Ai
Modelpublic

diffusers/LTX-2.3-Diffusers

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
4likes1.3kdownloads
README.md170 linesDownload Raw Back to root
1---2library_name: diffusers3pipeline_tag: image-to-video4base_model: Lightricks/LTX-2.35tags:6- image-to-video7- text-to-video8- video-to-video9- image-text-to-video10- audio-to-video11- text-to-audio12- video-to-audio13- audio-to-audio14- text-to-audio-video15- image-to-audio-video16- image-text-to-audio-video17- ltx-218- ltx-2-319- ltx-video20- ltxv21- lightricks22license: other23license_name: ltx-video-2-open-source-license24license_link: https://huggingface.co/Lightricks/LTX-2.3/blob/main/LICENSE25---26 27# LTX-2.3 (Diffusers)28 29Diffusers-format weights for [Lightricks/LTX-2.3](https://huggingface.co/Lightricks/LTX-2.3) — a DiT-based foundation model that jointly generates synchronized video and audio.30 31A distilled variant (8 steps, CFG=1) is available at [`diffusers/LTX-2.3-Distilled-Diffusers`](https://huggingface.co/diffusers/LTX-2.3-Distilled-Diffusers).32 33## Usage34 35Requires a recent build of `diffusers` with LTX-2 support:36 37```bash38pip install -U git+https://github.com/huggingface/diffusers39```40 41### Text-to-video + audio42 43```python44import torch45from diffusers import LTX2Pipeline46from diffusers.pipelines.ltx2.export_utils import encode_video47from diffusers.pipelines.ltx2.utils import DEFAULT_NEGATIVE_PROMPT48 49pipe = LTX2Pipeline.from_pretrained(50    "diffusers/LTX-2.3-Diffusers", torch_dtype=torch.bfloat1651)52pipe.enable_model_cpu_offload()53 54prompt = "A flowing river in a forest at golden hour, gentle wind in the leaves."55frame_rate = 24.056 57video, audio = pipe(58    prompt=prompt,59    negative_prompt=DEFAULT_NEGATIVE_PROMPT,60    width=768,61    height=512,62    num_frames=121,63    frame_rate=frame_rate,64    num_inference_steps=30,65    guidance_scale=3.0,66    output_type="np",67    return_dict=False,68)69 70encode_video(71    video[0],72    fps=frame_rate,73    audio=audio[0].float().cpu(),74    audio_sample_rate=pipe.vocoder.config.output_sampling_rate,75    output_path="ltx2_t2v.mp4",76)77```78 79### First-last-frame-to-video (FLF2V)80 81```python82import torch83from diffusers import LTX2ConditionPipeline84from diffusers.pipelines.ltx2.pipeline_ltx2_condition import LTX2VideoCondition85from diffusers.pipelines.ltx2.utils import DEFAULT_NEGATIVE_PROMPT86from diffusers.utils import load_image87 88pipe = LTX2ConditionPipeline.from_pretrained(89    "diffusers/LTX-2.3-Diffusers", torch_dtype=torch.bfloat1690)91pipe.enable_model_cpu_offload()92 93first_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/flf2v_input_first_frame.png")94last_image = load_image("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/flf2v_input_last_frame.png")95 96conditions = [97    LTX2VideoCondition(frames=first_image, index=0, strength=1.0),98    LTX2VideoCondition(frames=last_image, index=-1, strength=1.0),99]100 101prompt = "CG animation style, a small blue bird takes off from the ground, flapping its wings."102frame_rate = 24.0103 104video = pipe(105    conditions=conditions,106    prompt=prompt,107    negative_prompt=DEFAULT_NEGATIVE_PROMPT,108    width=768,109    height=512,110    num_frames=121,111    frame_rate=frame_rate,112    num_inference_steps=40,113    guidance_scale=4.0,114    output_type="np",115    return_dict=False,116)117```118 119### IC-LoRA (camera control)120 121```python122import torch123from diffusers import LTX2InContextPipeline124from diffusers.pipelines.ltx2.export_utils import encode_video125from diffusers.pipelines.ltx2.utils import DEFAULT_NEGATIVE_PROMPT126 127pipe = LTX2InContextPipeline.from_pretrained(128    "diffusers/LTX-2.3-Diffusers", torch_dtype=torch.bfloat16129)130pipe.enable_model_cpu_offload()131pipe.load_lora_weights(132    "Lightricks/LTX-2-19b-LoRA-Camera-Control-Dolly-In",133    adapter_name="ic_lora",134    weight_name="ltx-2-19b-lora-camera-control-dolly-in.safetensors",135)136pipe.set_adapters("ic_lora", 1.0)137 138prompt = "A flowing river in a forest"139frame_rate = 24.0140 141video, audio = pipe(142    prompt=prompt,143    negative_prompt=DEFAULT_NEGATIVE_PROMPT,144    width=768,145    height=512,146    num_frames=121,147    frame_rate=frame_rate,148    num_inference_steps=30,149    guidance_scale=3.0,150    output_type="np",151    return_dict=False,152)153 154encode_video(155    video[0],156    fps=frame_rate,157    audio=audio[0].float().cpu(),158    audio_sample_rate=pipe.vocoder.config.output_sampling_rate,159    output_path="ltx2_ic_lora.mp4",160)161```162 163## Notes164 165- `width` and `height` must be divisible by 32; `num_frames` must equal `8k + 1`.166- See the [Diffusers LTX-2 docs](https://huggingface.co/docs/diffusers/main/en/api/pipelines/ltx2) for multimodal guidance, prompt enhancement, and the upscaling/refinement pipeline.167 168## License169 170These weights are released under the [LTX Video 2 Open Source License](https://huggingface.co/Lightricks/LTX-2.3/blob/main/LICENSE).