Team Ai
Modelpublic

mlx-community/Lance-3B-Video-bf16

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
10likes488downloads
Model Card
๐Ÿ“‚ Part of the [Lance MLX collection](https://huggingface.co/collections/mlx-community/lance-mlx-6a0f3cd5648a74f8283fc8a4) on mlx-community.

Lance-3B-Video-bf16 (MLX, video specialist) โ€” ๐Ÿšง ALPHA

MLX port of ByteDance Intelligent Creation Lab's Lance โ€” the video-specialist Lance_3B_Video checkpoint, converted to bf16 for Apple Silicon. ~6.44 B LLM parameters + 669 M Qwen2.5-VL ViT bundled, with the 126,976-entry latent_pos_embed table needed for video-scale latent grids.

โš ๏ธ This is an alpha release. t2v is production-quality through 768ยฒร—25f (nlat=16,128) after the Phase 5m CFG-renorm fix (v0.5.2), verified across two prompts (panda surfing, bus + Big Ben). At nlat โ‰ฅ ~30k (768ยฒร—49f, 480ร—848ร—121f) Phase 5m partially closes the original "pure noise" failure to a milder "structured-but-degraded with mesh artifacts" failure โ€” the model attempts the scene but the VAE outputs colored geometric tiles overlaid on it. See Status below.
Lance is ByteDance's 3B-active unified multimodal model (paper, code, HF original). This is not Lance/LanceDB, the columnar data format.

Status

๐ŸŸก Alpha โ€” t2v is production-quality through 768ยฒร—25f (n_lat=16,128) after the Phase 5m fix; n_lat โ‰ฅ ~30k produces structured-but-degraded output with mesh artifacts (improvement over pre-fix pure noise, but not usable); understanding pipelines unvalidated.

CapabilityStatusNotes
t2v at 256ร—256 ร— โ‰ค17 frames๐ŸŸข ProductionRed panda surfing demo shows real temporal motion. ~33 s/clip on M5 Max.
t2v at 512ร—512 ร— 17 frames (n_lat โ‰ค 5,120)๐ŸŸข ProductionPainterly aesthetic (this checkpoint's training-time style).
t2v at 640ร—640 ร— 17 frames (n_lat โ‰ค 8,000)๐ŸŸข ProductionNew scale validated under Phase 5m.
t2v at 768ร—768 ร— โ‰ค13 frames (n_lat โ‰ค 9,216)๐ŸŸข ProductionPainterly; legacy baseline.
t2v at 768ร—768 ร— 17 frames (n_lat = 11,520)๐ŸŸข Production (Phase 5m)CFG-renorm fix in [v0.5.2](https://github.com/xocialize/lance-mlx/releases/tag/v0.5.2-phase5m-cfg-renorm) โ€” closes the silent quality regression. Pass `cfg_renorm_type="global"` to restore legacy default.
t2v at 768ร—768 ร— 25 frames (n_lat = 16,128)๐ŸŸข Production (Phase 5m+)Verified post-fix with clean diagnostic prompt (bus + Big Ben). Quality equivalent-or-better than 17f; the n_lat โ†’ quality relationship is stochastic (seed ร— scale), not a monotonic degradation curve.
t2v at 768ร—768 ร— 33-41 frames (n_lat = 21kโ€“26k)โ“ Untested with Phase 5mGap in the empirical sweep. Likely in-envelope based on 25f result but unverified.
t2v at 768ร—768 ร— 49 frames (n_lat = 29,952)โŒ Structured-but-degradedManual verification 2026-05-23: Phase 5m partially closes the pre-fix pure-noise collapse to a milder failure โ€” the model attempts the scene (Big Ben silhouette barely visible) but the VAE produces colored geometric mesh artifacts overlaid throughout. Numerical signature: final std=0.623 vs ~0.88 for clean runs โ€” channel renorm clamps too aggressively at late timesteps once n_lat reaches 30k, pushing latents outside the VAE's trained distribution. ~78 min wall-clock + 84.6 GB peak memory. Tracked as issue #1.
t2v at Lance reference (480ร—848 ร— 121f, n_lat โ‰ˆ 49 k)โŒ Same regime as aboveUntested directly at this exact dimension but expected to fall in the same degraded-mesh-artifact regime as 768ยฒร—49f.
x2t_video (video VQA / captioning)๐ŸŸก Implemented, not validatedPipeline lands in lance-mlx but hasn't been compared against Phase 0 oracle.
video_edit (instruction-based)๐ŸŸก Implemented, not validatedDirect fusion of t2v + image_edit. Will only be as good as t2v at the chosen scale.

For production-quality image tasks (t2i, imageedit, x2timage), use the sibling repo `mlx-community/Lance-3B-bf16` โ€” it's fully validated.

Hardware envelope (memory_mode, 2026-06-02)

The lance-mlx source repo's `memory_mode` knob (auto / parallel / relay) brings bf16 video generation within reach of 8โ€“16 GB Apple Silicon Macs:

RAMModet2v / video_editNotes
8โ€“16 GBrelay (auto-resolved)โœ… Long t2v clips fit (e.g. 256ยฒร—61f in ~8.7 GB)Single-shot per pipeline load โ€” re-prefill reloads the UND tower. Tiled spatial+temporal VAE decode keeps the decode peak roughly flat in frame count.
24 GB+parallel (auto-resolved)โœ… Reusable pipeline across callsAll components resident; classic behavior.

relay produces byte-identical output to parallel (frame MD5-verified on real Lance-3B-Video-bf16) โ€” it sheds the UND tower after prefill and frees the GEN tower before VAE decode, so peak memory โ‰ˆ heaviest single phase rather than the sum of all three. The tiled VAE decode (tile_vae=True, plus vae_temporal_tile for long clips) bounds the decode transient, so long t2v becomes loop-bound rather than decode-bound. Default auto resolves by mx.device_info()'s recommended working-set size with the split at ~18 GiB. The same envelope applies to the sibling `mlx-community/Lance-3B-bf16` for image tasks.

Within the n_lat โ‰ค 16,128 production envelope above, relay fits every validated configuration in 16 GB; the larger-than-envelope configs (n_lat โ‰ฅ ~30k) are still gated by the structured-mesh-artifact issue documented below, not by memory.

Lossless streaming VAE decode (lossless_decode, 2026-06-05)

generate(..., lossless_decode=True) (default since 2026-06-05) uses a bit-identical streaming VAE decode that's flat in frame count via temporal causal-cache streaming (Wan2.2 VAE is causal in time), composed with spatial halo-tile + crop. 50-case bit-identity test in `lance-mlx` with a negative control verifies the guarantee (max|ฮ”|=0 vs whole-sequence dec(z)). On real Lance-3B-Video-bf16 with ri_phys measurement (true OS-committed footprint, supersedes the older mx.get_peak_memory() quotes which under-report ~2ร—):

Configwhole `dec(z)`lossless streaminglossy blend16 GB
256ยฒ ร— 49f12.18 GB7.99 GBโ€”โœ… lossless (lighter + exact)
256ยฒ ร— 121f15.41 GB8.05 GB12.47 GBโœ… lossless (lighter + exact)
512ยฒ ร— 61fOOM12.64 GBOOM (>20 GB)โœ… lossless (only path that fits)
768ยฒ video (โ‰ฅ13f)OOM>~21 GB (lower bound)13.2โ€“16.3 GB (+~3 GB swap)lossy on 16 GB (lossless needs 24 GB+, predicted)

The temporal streaming win widens with clip length โ€” at 256ยฒ the lossless decode goes from 7.99 GB at 49f to 8.05 GB at 121f, essentially flat. Only 768ยฒ video needs the lossy fallback on 16 GB: pass lossless_decode=False to keep the trapezoidal-blend tiling (~1.5โ€“4.8 / 255 off the reference). For bit-identical 768ยฒ video, a 24 GB+ Mac is needed (predicted from the >~21 GB lower bound, not directly measured).

Full envelope, methodology, and raw numbers in the source repo's `LIMITS.md`.

There is no AWQ-INT4 quantized variant for video โ€” mlx-community/Lance-3B-AWQ-INT4 ships for x2t_image VQA only and is not applicable to t2v / videoedit / x2tvideo. For video on small Macs, bf16 + memory_mode=relay (+ lossless_decode=False for 768ยฒ video) is the path.

Why a separate "Video" checkpoint?

ByteDance ships two variants of Lance that differ in fine-tuning (NOT just latent_pos_embed size):

  • โ€”Lance_3B โ€” image specialist. Crystal-clear photorealistic t2i.
  • โ€”Lance_3B_Video โ€” video specialist. Same architecture, further fine-tuned on video data. Native aesthetic is painterly (verified by per-tensor diff: _moe_gen QK-norms differ by 0.5โ€“0.85 in 6+ layers; lm_head and embed_tokens are byte-identical).

This checkpoint also bundles the Qwen2.5-VL ViT for video-understanding tasks, with the larger 126,976-entry latent_pos_embed table that addresses video-resolution token grids.

Quickstart

Install from the lance-mlx source repo:

bash
git clone https://github.com/xocialize/lance-mlx
cd lance-mlx && uv sync

Download this checkpoint:

python
from huggingface_hub import snapshot_download
weights = snapshot_download("mlx-community/Lance-3B-Video-bf16")

Text-to-video (recommended scale)

python
from lance_mlx.pipeline.t2v import TextToVideoPipeline

pipe = TextToVideoPipeline.from_pretrained(
    lance_weights_dir=weights,
    vae_safetensors=f"{weights}/vae.safetensors",
)
frames = pipe.generate(
    "A red panda surfing on a sunny wave.",
    num_frames=16, height=256, width=256,
    num_steps=30, cfg_scale=4.0, seed=42,
)
# frames is np.ndarray of shape (T_decoded, H, W, 3) uint8

Encode to MP4 with imageio:

python
import imageio
with imageio.get_writer("out.mp4", fps=12, codec="libx264") as writer:
    for f in frames:
        writer.append_data(f)

Video understanding (alpha)

python
from lance_mlx.pipeline.understanding import UnderstandingPipeline

pipe = UnderstandingPipeline.from_pretrained(
    lance_weights_dir=weights,
    vit_safetensors=f"{weights}/vit.safetensors",
)
answer = pipe.generate_video(
    video="my_video.mp4",
    question="Describe what happens in this video.",
    num_sample_frames=16, target_h=224, target_w=224,
    max_new_tokens=256, prompt_style="lance",
)
print(answer)

โš ๏ธ Unvalidated against Phase 0 oracle. Treat answers as exploratory.

Phase 5m fix โ€” silent quality regression at n_lat โ‰ˆ 11,520 RESOLVED (v0.5.2)

The "global" CFG-renorm cap was computing a single scalar L2 over the entire velocity tensor. At higher n_lat (โ‰ˆ 2ร— the production baseline) the L2 sum spans roughly twice as many elements, so the same cap silently over-suppressed high-frequency detail โ€” composition + identity correct, but textures and sky gradients degraded.

Fix (default since v0.5.2): cfg_renorm_type="channel" computes per-channel L2 separately, so pathological channels clamp without dragging the aggregate signal down. Detail returns at high n_lat without regressing small scales.

Evidence (768ยฒร—17f, seed=43, cfg=4.0): "global" final std 0.907 (over-suppressed), "none" final std 1.112 (uncapped + clean), "channel" final std 0.900 (capped per-channel + visually matches "none"). V0 safety A/B at 768ยฒร—13f confirms no small-scale regression.

Pass cfg_renorm_type="global" to restore the legacy default.

Known issue: structured-but-degraded mesh artifacts at n_lat โ‰ฅ ~30k

Lance_3B_Video t2v pre-Phase-5m collapsed to pure random noise at very-high latent counts. Post-Phase-5m the failure mode is milder but still unusable: the model attempts the scene (recognizable silhouettes barely visible) but the VAE outputs colored geometric mesh tiles overlaid throughout.

Bisection on Phase 5m defaults (cfg_renorm_type="channel"):

T_frames   n_lat   result
       1   2,304   coherent  (same as t2i)
       5   4,608   coherent
       9   6,912   coherent (painterly)
      13   9,216   coherent (painterly, mild temporal drift)
      17  11,520   coherent โ† Phase 5m fix restored detail
      25  16,128   coherent โ† Phase 5m+ verified across two prompts
      33  21,120   untested
      41  26,304   untested
      49  29,952   structured-but-degraded โ† manual verification 2026-05-23

Numerical signature of the degraded regime: final std=0.623 (49f) vs ~0.88 (clean runs). Channel renorm clamps too aggressively at late timesteps once n_lat reaches ~30k, pushing latents outside the VAE's trained distribution. The mesh-tile pattern is the VAE's response to out-of-distribution latents โ€” not random noise but a low-rank geometric approximation.

Open candidates for a future Phase 5n / issue #1 fix:

  • โ€”Per-channel renorm threshold that scales with n_lat (currently constant)
  • โ€”Alternative late-timestep clamping (e.g. cfg_interval=[0.4, 1.0] to disable CFG entirely in the last steps)
  • โ€”Investigating whether VAE decoder can be retrained on Phase-5m-style latents (longer-term)

The bug does not affect:

  • โ€”Image tasks (use `mlx-community/Lance-3B-bf16`).
  • โ€”t2v through 768ยฒ ร— 25f with Phase 5m defaults.
  • โ€”The model checkpoint itself โ€” same weights produce coherent images at any resolution.

Tracked at github.com/xocialize/lance-mlx/issues/1.

Performance (M5 Max 128 GB)

TaskConfigurationWall-clock
t2v256ยฒ ร— 16f, 30 steps, CFG=4.0~33 s
t2v512ยฒ ร— 16f, 30 steps, CFG=4.0~60 s
t2v768ยฒ ร— 13f, 30 steps, CFG=4.0~145 s

CFG doubles the forward cost since cond + uncond run sequentially. KV cache for the text prefix is a Phase 5 follow-up.

Files in this repo

FileSizePurpose
model.safetensors12.87 GBLLM weights (1021 tensors, both UND + GEN towers, with 126,976-entry latentposembed)
vit.safetensors1.34 GBQwen2.5-VL ViT (semantic encoder for x2t_video)
vae.safetensors1.41 GBLance's bundled Wan2.2 VAE (also available standalone as `mlx-community/Wan2.2-VAE-Lance-bf16`)
config.jsonโ€“Qwen2_5_VLForConditionalGeneration config
conversion_report.jsonโ€“Provenance
tokenizer.json / vocab.jsonโ€“Qwen2.5-VL vocabulary

Provenance

Source: bytedance-research/Lance/Lance_3B_Video/model.safetensors (1411 tensors including bundled ViT; 6.437 B LLM + 0.669 B ViT params). Converted via `scripts/02_convert.py`. The bundled ViT is extracted to a sibling vit.safetensors with the vit_model. prefix stripped, matching the layout convention of the image-specialist repo.

Limitations + caveats

  • โ€”Aesthetic is painterly by design. Lance3BVideo was further fine-tuned on video data; the native style is intentionally painterly, not photorealistic. Lance_3B (image specialist) is the crystal-photo checkpoint.
  • โ€”Pending-verification regime at nlat โ‰ฅ ~30k (see [Known issue](#known-issue-pending-verification-regime-at-nlat--30k)). Phase 5m fixed the silent quality regression at intermediate n_lat (verified through 16,128 with channel renorm).
  • โ€”No streaming or batched generation.
  • โ€”English + Chinese prompts. Other languages are out of distribution.

License

This MLX port: Apache 2.0.

Underlying weights:

  • โ€”Lance: Apache 2.0 (ByteDance Intelligent Creation Lab).
  • โ€”Wan2.2 VAE: Apache 2.0 (Alibaba).
  • โ€”Qwen2.5-VL: Apache 2.0 (Alibaba).

See `NOTICE` for attribution.

Citation

bibtex
@article{fu2026lance,
  title={Lance: Unified Multimodal Modeling by Multi-Task Synergy},
  author={Fu, Fengyi and Huang, Mengqi and Wu, Shaojin and others},
  journal={arXiv preprint arXiv:2605.18678},
  year={2026}
}

Links