Team Ai
Modelpublic

AsadIsmail/CogVideoX-5b-ternary

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes15downloads
Model Card

CogVideoX-5B — Ternary Quantized (tritplane3)

Storage format note. This checkpoint is stored dequantized to FP16. Loaded with stock diffusers it runs like the base model — the memory/compute benefits of ternary are not realized in this format. It demonstrates that ternary PTQ preserves generation quality. The packed 2-bit ternary runtime (`ternary-quant`) currently targets transformers LLMs/VLMs, not diffusers video pipelines, so there is no accelerated ternary inference path for this model yet. Packed weights and a diffusers runtime are on the roadmap.

Ternary-quantized version of zai-org/CogVideoX-5b produced with ternary-quant using component-aware tritplane3 quantization applied to the Diffusion Transformer (DiT) backbone.

This is an experimental diffusers-compatible artifact. It is not a benchmarked replacement for FP8, int8, or other production video quantization paths.

Model Specifications

PropertyValue
Base Modelzai-org/CogVideoX-5b
ArchitectureDiffusion Transformer (CogVideoXTransformer3DModel)
Transformer Params5.57B
Quantizationtritplane3 (3-plane progressive ternary)
Components Quantized341 linear layers in the DiT
Text Encoder (T5)FP16 (preserved)
VAE (3D causal)FP16 (preserved)
LicenseApache 2.0

Size & Compression

MethodTransformer SizeBits/WeightCompression
FP16 (original)11.14 GB161.0×
Ternary tritplane3 (theoretical, packed)~5.57 GB~82.0×
FP16 (as stored in this repo)11.14 GB161.0× on disk

Honest note: Weights have ternary precision but are stored in FP16 format for drop-in compatibility with the standard diffusers pipeline. For actual 2× disk compression, weights would need packed tritplane format (requires custom inference wrapper).

Memory Requirements (Inference)

DevicePeak MemoryRecommendation
Apple Silicon MPS (bfloat16)~24 GB unifiedM2 Pro 32GB+ or M4 Pro 24GB+
NVIDIA CUDA (bfloat16)~20 GB VRAMRTX 4090 / A6000
CPUNot recommendedToo slow

Quickstart

bash
pip install diffusers transformers accelerate tiktoken sentencepiece protobuf imageio imageio-ffmpeg
python
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import export_to_video

pipe = DiffusionPipeline.from_pretrained(
    "AsadIsmail/CogVideoX-5b-ternary",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=True,
)

device = "mps"  # or "cuda"
if device == "mps":
    for attr in ("alphas_cumprod", "betas", "alphas", "sigmas"):
        val = getattr(pipe.scheduler, attr, None)
        if torch.is_tensor(val) and val.dtype == torch.float64:
            setattr(pipe.scheduler, attr, val.float())

pipe.to(device)
pipe.enable_attention_slicing()

result = pipe(
    prompt="a panda playing bass guitar on stage",
    num_frames=9,
    num_inference_steps=25,
    guidance_scale=6.0,
    height=480, width=720,
    generator=torch.Generator(device=device).manual_seed(42),
)

export_to_video(result.frames[0], "output.mp4", fps=8)

Collection

Part of ternary-models.

GitHub: github.com/Asad-Ismail/ternary-models | Library: github.com/Asad-Ismail/ternary-quant