mlx-community/Lance-3B-Video-bf16
๐ Part of the [Lance MLX collection](https://huggingface.co/collections/mlx-community/lance-mlx-6a0f3cd5648a74f8283fc8a4) on mlx-community.
Lance-3B-Video-bf16 (MLX, video specialist) โ ๐ง ALPHA
MLX port of ByteDance Intelligent Creation Lab's Lance โ the video-specialist Lance_3B_Video checkpoint, converted to bf16 for Apple Silicon. ~6.44 B LLM parameters + 669 M Qwen2.5-VL ViT bundled, with the 126,976-entry latent_pos_embed table needed for video-scale latent grids.
โ ๏ธ This is an alpha release. t2v is production-quality through 768ยฒร25f (nlat=16,128) after the Phase 5m CFG-renorm fix (v0.5.2), verified across two prompts (panda surfing, bus + Big Ben). At nlat โฅ ~30k (768ยฒร49f, 480ร848ร121f) Phase 5m partially closes the original "pure noise" failure to a milder "structured-but-degraded with mesh artifacts" failure โ the model attempts the scene but the VAE outputs colored geometric tiles overlaid on it. See Status below.
Lance is ByteDance's 3B-active unified multimodal model (paper, code, HF original). This is not Lance/LanceDB, the columnar data format.
Status
๐ก Alpha โ t2v is production-quality through 768ยฒร25f (n_lat=16,128) after the Phase 5m fix; n_lat โฅ ~30k produces structured-but-degraded output with mesh artifacts (improvement over pre-fix pure noise, but not usable); understanding pipelines unvalidated.
For production-quality image tasks (t2i, imageedit, x2timage), use the sibling repo `mlx-community/Lance-3B-bf16` โ it's fully validated.
Hardware envelope (memory_mode, 2026-06-02)
The lance-mlx source repo's `memory_mode` knob (auto / parallel / relay) brings bf16 video generation within reach of 8โ16 GB Apple Silicon Macs:
relay produces byte-identical output to parallel (frame MD5-verified on real Lance-3B-Video-bf16) โ it sheds the UND tower after prefill and frees the GEN tower before VAE decode, so peak memory โ heaviest single phase rather than the sum of all three. The tiled VAE decode (tile_vae=True, plus vae_temporal_tile for long clips) bounds the decode transient, so long t2v becomes loop-bound rather than decode-bound. Default auto resolves by mx.device_info()'s recommended working-set size with the split at ~18 GiB. The same envelope applies to the sibling `mlx-community/Lance-3B-bf16` for image tasks.
Within the n_lat โค 16,128 production envelope above, relay fits every validated configuration in 16 GB; the larger-than-envelope configs (n_lat โฅ ~30k) are still gated by the structured-mesh-artifact issue documented below, not by memory.
Lossless streaming VAE decode (lossless_decode, 2026-06-05)
generate(..., lossless_decode=True) (default since 2026-06-05) uses a bit-identical streaming VAE decode that's flat in frame count via temporal causal-cache streaming (Wan2.2 VAE is causal in time), composed with spatial halo-tile + crop. 50-case bit-identity test in `lance-mlx` with a negative control verifies the guarantee (max|ฮ|=0 vs whole-sequence dec(z)). On real Lance-3B-Video-bf16 with ri_phys measurement (true OS-committed footprint, supersedes the older mx.get_peak_memory() quotes which under-report ~2ร):
The temporal streaming win widens with clip length โ at 256ยฒ the lossless decode goes from 7.99 GB at 49f to 8.05 GB at 121f, essentially flat. Only 768ยฒ video needs the lossy fallback on 16 GB: pass lossless_decode=False to keep the trapezoidal-blend tiling (~1.5โ4.8 / 255 off the reference). For bit-identical 768ยฒ video, a 24 GB+ Mac is needed (predicted from the >~21 GB lower bound, not directly measured).
Full envelope, methodology, and raw numbers in the source repo's `LIMITS.md`.
There is no AWQ-INT4 quantized variant for video โ mlx-community/Lance-3B-AWQ-INT4 ships for x2t_image VQA only and is not applicable to t2v / videoedit / x2tvideo. For video on small Macs, bf16 + memory_mode=relay (+ lossless_decode=False for 768ยฒ video) is the path.
Why a separate "Video" checkpoint?
ByteDance ships two variants of Lance that differ in fine-tuning (NOT just latent_pos_embed size):
Lance_3Bโ image specialist. Crystal-clear photorealistic t2i.Lance_3B_Videoโ video specialist. Same architecture, further fine-tuned on video data. Native aesthetic is painterly (verified by per-tensor diff:_moe_genQK-norms differ by 0.5โ0.85 in 6+ layers;lm_headandembed_tokensare byte-identical).
This checkpoint also bundles the Qwen2.5-VL ViT for video-understanding tasks, with the larger 126,976-entry latent_pos_embed table that addresses video-resolution token grids.
Quickstart
Install from the lance-mlx source repo:
git clone https://github.com/xocialize/lance-mlx
cd lance-mlx && uv syncDownload this checkpoint:
from huggingface_hub import snapshot_download
weights = snapshot_download("mlx-community/Lance-3B-Video-bf16")Text-to-video (recommended scale)
from lance_mlx.pipeline.t2v import TextToVideoPipeline
pipe = TextToVideoPipeline.from_pretrained(
lance_weights_dir=weights,
vae_safetensors=f"{weights}/vae.safetensors",
)
frames = pipe.generate(
"A red panda surfing on a sunny wave.",
num_frames=16, height=256, width=256,
num_steps=30, cfg_scale=4.0, seed=42,
)
# frames is np.ndarray of shape (T_decoded, H, W, 3) uint8Encode to MP4 with imageio:
import imageio
with imageio.get_writer("out.mp4", fps=12, codec="libx264") as writer:
for f in frames:
writer.append_data(f)Video understanding (alpha)
from lance_mlx.pipeline.understanding import UnderstandingPipeline
pipe = UnderstandingPipeline.from_pretrained(
lance_weights_dir=weights,
vit_safetensors=f"{weights}/vit.safetensors",
)
answer = pipe.generate_video(
video="my_video.mp4",
question="Describe what happens in this video.",
num_sample_frames=16, target_h=224, target_w=224,
max_new_tokens=256, prompt_style="lance",
)
print(answer)โ ๏ธ Unvalidated against Phase 0 oracle. Treat answers as exploratory.
Phase 5m fix โ silent quality regression at n_lat โ 11,520 RESOLVED (v0.5.2)
The "global" CFG-renorm cap was computing a single scalar L2 over the entire velocity tensor. At higher n_lat (โ 2ร the production baseline) the L2 sum spans roughly twice as many elements, so the same cap silently over-suppressed high-frequency detail โ composition + identity correct, but textures and sky gradients degraded.
Fix (default since v0.5.2): cfg_renorm_type="channel" computes per-channel L2 separately, so pathological channels clamp without dragging the aggregate signal down. Detail returns at high n_lat without regressing small scales.
Evidence (768ยฒร17f, seed=43, cfg=4.0): "global" final std 0.907 (over-suppressed), "none" final std 1.112 (uncapped + clean), "channel" final std 0.900 (capped per-channel + visually matches "none"). V0 safety A/B at 768ยฒร13f confirms no small-scale regression.
Pass cfg_renorm_type="global" to restore the legacy default.
Known issue: structured-but-degraded mesh artifacts at n_lat โฅ ~30k
Lance_3B_Video t2v pre-Phase-5m collapsed to pure random noise at very-high latent counts. Post-Phase-5m the failure mode is milder but still unusable: the model attempts the scene (recognizable silhouettes barely visible) but the VAE outputs colored geometric mesh tiles overlaid throughout.
Bisection on Phase 5m defaults (cfg_renorm_type="channel"):
T_frames n_lat result
1 2,304 coherent (same as t2i)
5 4,608 coherent
9 6,912 coherent (painterly)
13 9,216 coherent (painterly, mild temporal drift)
17 11,520 coherent โ Phase 5m fix restored detail
25 16,128 coherent โ Phase 5m+ verified across two prompts
33 21,120 untested
41 26,304 untested
49 29,952 structured-but-degraded โ manual verification 2026-05-23Numerical signature of the degraded regime: final std=0.623 (49f) vs ~0.88 (clean runs). Channel renorm clamps too aggressively at late timesteps once n_lat reaches ~30k, pushing latents outside the VAE's trained distribution. The mesh-tile pattern is the VAE's response to out-of-distribution latents โ not random noise but a low-rank geometric approximation.
Open candidates for a future Phase 5n / issue #1 fix:
- Per-channel renorm threshold that scales with n_lat (currently constant)
- Alternative late-timestep clamping (e.g. cfg_interval=[0.4, 1.0] to disable CFG entirely in the last steps)
- Investigating whether VAE decoder can be retrained on Phase-5m-style latents (longer-term)
The bug does not affect:
- Image tasks (use `mlx-community/Lance-3B-bf16`).
- t2v through 768ยฒ ร 25f with Phase 5m defaults.
- The model checkpoint itself โ same weights produce coherent images at any resolution.
Tracked at github.com/xocialize/lance-mlx/issues/1.
Performance (M5 Max 128 GB)
CFG doubles the forward cost since cond + uncond run sequentially. KV cache for the text prefix is a Phase 5 follow-up.
Files in this repo
Provenance
Source: bytedance-research/Lance/Lance_3B_Video/model.safetensors (1411 tensors including bundled ViT; 6.437 B LLM + 0.669 B ViT params). Converted via `scripts/02_convert.py`. The bundled ViT is extracted to a sibling vit.safetensors with the vit_model. prefix stripped, matching the layout convention of the image-specialist repo.
Limitations + caveats
- Aesthetic is painterly by design. Lance3BVideo was further fine-tuned on video data; the native style is intentionally painterly, not photorealistic. Lance_3B (image specialist) is the crystal-photo checkpoint.
- Pending-verification regime at nlat โฅ ~30k (see [Known issue](#known-issue-pending-verification-regime-at-nlat--30k)). Phase 5m fixed the silent quality regression at intermediate n_lat (verified through 16,128 with channel renorm).
- No streaming or batched generation.
- English + Chinese prompts. Other languages are out of distribution.
License
This MLX port: Apache 2.0.
Underlying weights:
- Lance: Apache 2.0 (ByteDance Intelligent Creation Lab).
- Wan2.2 VAE: Apache 2.0 (Alibaba).
- Qwen2.5-VL: Apache 2.0 (Alibaba).
See `NOTICE` for attribution.
Citation
@article{fu2026lance,
title={Lance: Unified Multimodal Modeling by Multi-Task Synergy},
author={Fu, Fengyi and Huang, Mengqi and Wu, Shaojin and others},
journal={arXiv preprint arXiv:2605.18678},
year={2026}
}Links
- MLX port code + phase notes: `github.com/xocialize/lance-mlx`
- Open issue (t2v scale collapse): #1
- Original PyTorch model: `bytedance-research/Lance`
- Image specialist (production): `mlx-community/Lance-3B-bf16`
- Wan2.2 VAE (standalone): `mlx-community/Wan2.2-VAE-Lance-bf16`
