MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
Yipeng Du, Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou, Xiang Li, Jian Yang, Zhenheng Yang, Ying Tai
Abstract
Despite advancements in Multimodal Large Language Models (MLLMs), their proficiency in fine-grained video motion understanding remains critically limited. They often lack inter-frame differencing and tend to average or ignore subtle visual cues. Furthermore, while visual prompting has shown potential in static images, its application to videos' temporal complexities, particularly for fine-grained motion understanding, remains largely unexplored. We investigate whether inherent capability can be unlocked to boost MLLMs' motion perception and enable distinct visual signatures tailored to decouple object and camera motion cues. In this study, we introduce , a novel zero-shot method pioneering object-centric visual spotlight and motion blur as visual prompts to effectively improve fine-grained motion understanding without training. To convert this into valuable data assets, we curated , the first large-scale dataset for fine-grained video motion understanding, with hierarchical annotations including SFT and preference data, video clips and QAs. Experiments show achieves state-of-the-art open-source performance and competitiveness with commercial models. Using , we fine-tuned on Qwen2.5VL-7B, which attains 48.3% overall accuracy on FAVOR-Bench that is comparable to Qwen2.5VL-72B's 48.1%. In summary, we present a novel zero-shot method and a large-scale, high-quality dataset specifically for fine-grained motion understanding. All the code and annotations will be publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Video Panels for Long Video UnderstandingLars Doorenbos, Federico Spurio, Juergen GallCVPR 2026 · 6 citations
- Building a Precise Video Language with Human–AI OversightZhiqiu Lin, Siyuan Cen, Chancharik Mitra, Isaac Li et al.CVPR 2026 · 3 citations
- MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language ModelsYifan Xu, Chao Zhang, Ruifei Ma, Fei Gao et al.CVPR 2026
Builds on11
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
- VideoChat-Flash: Hierarchical Compression for Long-Context Video ModelingXinhao Li, Yi Wang, Jiashuo Yu, Xiangyu Zeng et al.ICLR 2026 · 172 citations
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang et al.NeurIPS 2023 · 129 citations
- Intern VL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic TasksZhe Chen, Jiannan Wu, Wenhai Wang, Weijie Su et al.CVPR 2024
Related papers
- TIME: Temporal-Sensitive Multi-Dimensional Instruction Tuning and Robust Benchmarking for Video-LLMsYunxiao Wang, Meng Liu, Wenqi Liu, Xuemeng Song et al.AAAI 2026 · 1 citation
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained UnderstandingYangliu Hu, Zikai Song, Na Feng, Yawei Luo et al.CVPR 2025
- TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video UnderstandingShuhuai Ren, Linli Yao, Shicheng Li, Xu Sun et al.CVPR 2024 · 83 citations
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language ModelsYifang Xu, Yunzhuo Sun, Benxiang Zhai, Ming Li et al.AAAI 2025 · 17 citations
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye et al.ICML 2024 · 121 citations
