MotionInsight

MotionInsight: Diagnosing Object Motion
Deficiencies in Generated Videos

EMNLP 2026 Findings

Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang,
Jingqi Tong, Junlin Li, Li Zhang, Shijie Zhao, Tianfan Xue

Code · arXiv · Quick start · Citation

MotionInsight-8B

MotionInsight diagnoses object motion in generated videos using RGB frames, object tracks, and camera motion. It produces reasoning and three scores from 1 (poor) to 5 (excellent):

Dimension Output field
Object consistency structural_stability
Motion continuity motion_coherence
Physical plausibility physical_plausibility

The model is based on Qwen3-VL-8B-Instruct and uses four BF16 Safetensors shards (approximately 17.7 GB).

Quick start

Use Python 3.10 and the custom Qwen3-VL patch from the code repository:

git clone --recurse-submodules https://github.com/JohnZhan2023/MotionInsight.git
cd MotionInsight
python3.10 -m venv .venv-model
source .venv-model/bin/activate
python -m pip install --upgrade pip
python -m pip install huggingface-hub==0.36.2
hf download JohnZhan/MotionInsight-8B --local-dir checkpoints/MotionInsight-8B
python -m pip install torch==2.11.0 torchvision==0.26.0 \
  --index-url https://download.pytorch.org/whl/cu126
python -m pip install -r checkpoints/MotionInsight-8B/requirements/inference.txt
python scripts/patch_transformers.py

After extracting motion features, run:

python inference.py \
  --model_name_or_path checkpoints/MotionInsight-8B \
  --video_path /path/to/video.mp4 \
  --object_motion_path /path/to/object_motion.pt \
  --camera_motion_path /path/to/camera_motion.pt \
  --target "tennis ball" \
  --output_path outputs/prediction.jsonl

Example answer format:

<thinking>Diagnostic reasoning about the target object's motion.</thinking>
<answer>{"structural_stability": 3.0, "physical_plausibility": 4.5, "motion_coherence": 3.5}</answer>

GRPO

The code repository includes a GRPO fine-tuning script:

python -m pip install -r checkpoints/MotionInsight-8B/requirements/train.txt
ATTN_IMPLEMENTATION=sdpa bash training/run_grpo.sh \
  checkpoints/MotionInsight-8B /path/to/train.jsonl outputs/motioninsight-grpo

See the repository README for the input format and preprocessing setup.

License

Model weights use Apache-2.0. Preprocessing dependencies retain their own licenses; see NOTICE.

Citation

@misc{zhan2026motioninsightdiagnosingobjectmotion,
  title         = {MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos},
  author        = {Jiahao Zhan and Yongrui Ma and Qunliang Xing and Xuanyu Zhang and Jingqi Tong and Junlin Li and Li zhang and Shijie Zhao and Tianfan Xue},
  year          = {2026},
  eprint        = {2609.37030},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2609.37030},
}
Downloads last month
74
Safetensors
Model size
832k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JohnZhan/MotionInsight-8B

Finetuned
(623)
this model

Paper for JohnZhan/MotionInsight-8B