Team Ai
Modelpublic

kagakouko/Spatial-Interactor-Qwen3-VL-4B

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
2likes199downloads
Model Card

<p align="center"> <img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/icon.png" width="100" alt="Spatial-Interactor"> </p>

<h1 align="center">Spatial-Interactor Qwen3-VL-4B</h1>

<p align="center"> <strong>Learning Spatial Reasoning through Interaction with the Observable Physical World</strong> </p>

<p align="center"> <a href="https://zju-omniai.github.io/Spatial-Interactor/"><img src="https://img.shields.io/badge/Project-Page-A56F59?style=flat-square&amp;labelColor=54534D" alt="Project page"></a> <a href="https://zju-omniai.github.io/Spatial-Interactor/assets/paper.pdf?v=20260917"><img src="https://img.shields.io/badge/Paper-PDF-9B8255?style=flat-square&amp;labelColor=54534D" alt="Paper PDF"></a> <a href="https://github.com/ZJU-OmniAI/Spatial-Interactor"><img src="https://img.shields.io/badge/Code-GitHub-738363?style=flat-square&amp;labelColor=54534D" alt="Code"></a> <a href="https://huggingface.co/datasets/kagakouko/LSI-108K"><img src="https://img.shields.io/badge/LSI--108K-Dataset-887A9A?style=flat-square&amp;labelColor=54534D" alt="Dataset"></a> </p>

This is the full-parameter BF16 Spatial-Interactor checkpoint based on Qwen/Qwen3-VL-4B-Instruct. It learns local world-state and ego-motion transitions through supervised fine-tuning, then uses On-Policy Distillation (OPD) to integrate successive transitions over long trajectories.

The privileged transition trace is used only during training. At inference, this checkpoint takes the same image/video and question inputs as its base model, with no extra trace, reward model, or teacher branch.

Overview

<p align="center"> <img src="https://raw.githubusercontent.com/ZJU-OmniAI/Spatial-Interactor/main/assets/readme/overview.webp?v=20260917" width="100%" alt="Spatial-Interactor overview: interaction trajectories, three-level curriculum, SFT and OPD, and spatial reasoning results"> </p>

<video src="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B/resolve/main/assets/presentation/spatial-interactor-intro-en.mp4?v=20260924-faithful" controls autoplay muted loop playsinline preload="metadata" poster="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B/resolve/main/assets/presentation/spatial-interactor-intro-en-poster.webp?v=20260924-faithful" width="100%"></video>

Presentation

<video src="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B/resolve/main/assets/presentation/spatial-interactor-presentation-en.mp4?v=20260924" controls autoplay muted loop playsinline preload="metadata" poster="https://huggingface.co/kagakouko/Spatial-Interactor-Qwen3-VL-4B/resolve/main/assets/presentation/presentation-en-poster.webp?v=20260924" width="100%"></video>

Load

python
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "kagakouko/Spatial-Interactor-Qwen3-VL-4B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto",
)

Use the base model's image/video input format. No privileged trace or additional teacher is needed for inference. Weights, tokenizer, processor, and chat template are included. See the training guide for SFT and OPD.

Citation

For citation, use the project BibTeX.

License

This checkpoint is released under Apache-2.0, following the base model license. Users must also comply with licenses and terms governing input datasets and media.