Team Ai
Apppublic

hugging-apps/vega-3d-spatial-reasoning

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes
App README

VEGA-3D · Spatial Reasoning

Interactive demo of **VEGA-3D**, a spatial-reasoning multimodal LLM from the paper *Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding* (ECCV 2026).

VEGA-3D augments a Qwen2.5-VL backbone with implicit 3D priors extracted from a frozen Wan2.1-T2V-1.3B video diffusion model. Intermediate spatiotemporal features from the diffusion model are fused into the VLM via token-level gated fusion, giving it stronger geometric and spatial understanding of indoor scenes.

Upload a scene image (or a short scene video) and ask a spatial-reasoning question.