JohnZhan/MotionInsight-8B
<p align="center"> <img src="assets/motioninsight-header.png" alt="MotionInsight" width="100%"> </p>
<h1 align="center">MotionInsight: Diagnosing Object Motion<br>Deficiencies in Generated Videos</h1> <p align="center"><strong>EMNLP 2026 Findings</strong></p> <p align="center"> Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang,<br> Jingqi Tong, Junlin Li, Li Zhang, Shijie Zhao, Tianfan Xue </p> <p align="center"> <a href="https://github.com/JohnZhan2023/MotionInsight">Code</a> · <a href="https://arxiv.org/abs/2609.37030">arXiv</a> · <a href="#quick-start">Quick start</a> · <a href="#citation">Citation</a> </p>
MotionInsight-8B
MotionInsight diagnoses object motion in generated videos using RGB frames, object tracks, and camera motion. It produces reasoning and three scores from 1 (poor) to 5 (excellent):
The model is based on Qwen3-VL-8B-Instruct and uses four BF16 Safetensors shards (approximately 17.7 GB).
Quick start
Use Python 3.10 and the custom Qwen3-VL patch from the code repository:
git clone --recurse-submodules https://github.com/JohnZhan2023/MotionInsight.git
cd MotionInsight
python3.10 -m venv .venv-model
source .venv-model/bin/activate
python -m pip install --upgrade pip
python -m pip install huggingface-hub==0.36.2
hf download JohnZhan/MotionInsight-8B --local-dir checkpoints/MotionInsight-8B
python -m pip install torch==2.11.0 torchvision==0.26.0 \
--index-url https://download.pytorch.org/whl/cu126
python -m pip install -r checkpoints/MotionInsight-8B/requirements/inference.txt
python scripts/patch_transformers.pyAfter extracting motion features, run:
python inference.py \
--model_name_or_path checkpoints/MotionInsight-8B \
--video_path /path/to/video.mp4 \
--object_motion_path /path/to/object_motion.pt \
--camera_motion_path /path/to/camera_motion.pt \
--target "tennis ball" \
--output_path outputs/prediction.jsonlExample answer format:
<thinking>Diagnostic reasoning about the target object's motion.</thinking>
<answer>{"structural_stability": 3.0, "physical_plausibility": 4.5, "motion_coherence": 3.5}</answer>GRPO
The code repository includes a GRPO fine-tuning script:
python -m pip install -r checkpoints/MotionInsight-8B/requirements/train.txt
ATTN_IMPLEMENTATION=sdpa bash training/run_grpo.sh \
checkpoints/MotionInsight-8B /path/to/train.jsonl outputs/motioninsight-grpoSee the repository README for the input format and preprocessing setup.
License
Model weights use Apache-2.0. Preprocessing dependencies retain their own licenses; see NOTICE.
Citation
@misc{zhan2026motioninsightdiagnosingobjectmotion,
title = {MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos},
author = {Jiahao Zhan and Yongrui Ma and Qunliang Xing and Xuanyu Zhang and Jingqi Tong and Junlin Li and Li zhang and Shijie Zhao and Tianfan Xue},
year = {2026},
eprint = {2609.37030},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2609.37030},
}