Team Ai
Modelpublic

JohnZhan/MotionInsight-8B

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
1likes71downloads
Model Card

<p align="center"> <img src="assets/motioninsight-header.png" alt="MotionInsight" width="100%"> </p>

<h1 align="center">MotionInsight: Diagnosing Object Motion<br>Deficiencies in Generated Videos</h1> <p align="center"><strong>EMNLP 2026 Findings</strong></p> <p align="center"> Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang,<br> Jingqi Tong, Junlin Li, Li Zhang, Shijie Zhao, Tianfan Xue </p> <p align="center"> <a href="https://github.com/JohnZhan2023/MotionInsight">Code</a> · <a href="https://arxiv.org/abs/2609.37030">arXiv</a> · <a href="#quick-start">Quick start</a> · <a href="#citation">Citation</a> </p>

MotionInsight-8B

MotionInsight diagnoses object motion in generated videos using RGB frames, object tracks, and camera motion. It produces reasoning and three scores from 1 (poor) to 5 (excellent):

DimensionOutput field
Object consistencystructural_stability
Motion continuitymotion_coherence
Physical plausibilityphysical_plausibility

The model is based on Qwen3-VL-8B-Instruct and uses four BF16 Safetensors shards (approximately 17.7 GB).

Quick start

Use Python 3.10 and the custom Qwen3-VL patch from the code repository:

bash
git clone --recurse-submodules https://github.com/JohnZhan2023/MotionInsight.git
cd MotionInsight
python3.10 -m venv .venv-model
source .venv-model/bin/activate
python -m pip install --upgrade pip
python -m pip install huggingface-hub==0.36.2
hf download JohnZhan/MotionInsight-8B --local-dir checkpoints/MotionInsight-8B
python -m pip install torch==2.11.0 torchvision==0.26.0 \
  --index-url https://download.pytorch.org/whl/cu126
python -m pip install -r checkpoints/MotionInsight-8B/requirements/inference.txt
python scripts/patch_transformers.py

After extracting motion features, run:

bash
python inference.py \
  --model_name_or_path checkpoints/MotionInsight-8B \
  --video_path /path/to/video.mp4 \
  --object_motion_path /path/to/object_motion.pt \
  --camera_motion_path /path/to/camera_motion.pt \
  --target "tennis ball" \
  --output_path outputs/prediction.jsonl

Example answer format:

text
<thinking>Diagnostic reasoning about the target object's motion.</thinking>
<answer>{"structural_stability": 3.0, "physical_plausibility": 4.5, "motion_coherence": 3.5}</answer>

GRPO

The code repository includes a GRPO fine-tuning script:

bash
python -m pip install -r checkpoints/MotionInsight-8B/requirements/train.txt
ATTN_IMPLEMENTATION=sdpa bash training/run_grpo.sh \
  checkpoints/MotionInsight-8B /path/to/train.jsonl outputs/motioninsight-grpo

See the repository README for the input format and preprocessing setup.

License

Model weights use Apache-2.0. Preprocessing dependencies retain their own licenses; see NOTICE.

Citation

bibtex
@misc{zhan2026motioninsightdiagnosingobjectmotion,
  title         = {MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos},
  author        = {Jiahao Zhan and Yongrui Ma and Qunliang Xing and Xuanyu Zhang and Jingqi Tong and Junlin Li and Li zhang and Shijie Zhao and Tianfan Xue},
  year          = {2026},
  eprint        = {2609.37030},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.37030},
}