hugging-apps/vega-3d-spatial-reasoning
0
VEGA-3D · Spatial Reasoning
Interactive demo of **VEGA-3D**, a spatial-reasoning multimodal LLM from the paper *Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding* (ECCV 2026).
VEGA-3D augments a Qwen2.5-VL backbone with implicit 3D priors extracted from a frozen Wan2.1-T2V-1.3B video diffusion model. Intermediate spatiotemporal features from the diffusion model are fused into the VLM via token-level gated fusion, giving it stronger geometric and spatial understanding of indoor scenes.
Upload a scene image (or a short scene video) and ask a spatial-reasoning question.
