Segment Any Motion in Videos
Nan Huang, Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer, Shanghang Zhang, Angjoo Kanazawa, Qianqian Wang
Abstract
Moving object segmentation is a crucial task for achieving a high-level understanding of visual scenes and has numerous downstream applications. Humans can effortlessly segment moving objects in videos. Previous work has largely relied on optical flow to provide motion cues; however, this approach often results in imperfect predictions due to challenges such as partial motion, complex deformations, motion blur and background distractions. We propose a novel approach for moving object segmentation that combines long-range trajectory motion cues with DINO-based semantic features and leverages SAM2 for pixel-level mask densification through an iterative prompting strategy. Our model employs Spatio-Temporal Trajectory Attention and Motion-Semantic Decoupled Embedding to prioritize motion while integrating semantic support. Extensive testing on diverse datasets demonstrates state-of-the-art performance, excelling in challenging scenarios and fine-grained segmentation of multiple objects. Our code is available at https://motion-seg.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- MoVieS: Motion-Aware 4D Dynamic View Synthesis in One SecondChenguo Lin, Yuchen Lin, Panwang Pan, Yifan Yu et al.CVPR 2026 · 38 citations
- Shape of Motion: 4D Reconstruction From a Single VideoQianqian Wang, Vickie Ye, Hang Gao, Weijia Zeng et al.ICCV 2025 · 29 citations
- SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video GenerationGuiyu Zhang, Yabo Chen, Xunzhi Xiang, Junchao Huang et al.CVPR 2026 · 8 citations
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsJiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li et al.NeurIPS 2025 · 7 citations
- AnthroTAP: Learning Point Tracking with Real-World MotionInès Hyeonsu Kim, Seokju Cho, Jahyeok Koo, Junghyun Park et al.CVPR 2026 · 5 citations
Builds on18
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- EmerNeRF: Emergent Spatial-Temporal Scene Decomposition via Self-SupervisionJiawei Yang, Boris Ivanovic, Or Litany, Xinshuo Weng et al.ICLR 2024 · 225 citations
- Motion-Attentive Transition for Zero-Shot Video Object SegmentationTianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao et al.AAAI 2020 · 210 citations
Related papers
- SAM2MOT: A Novel Paradigm of Multi-Object Tracking by SegmentationJunjie Jiang, Zelin Wang, Manqi Zhao, Yin Li et al.AAAI 2026 · 19 citations
- Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired DataHong Zhang, Yixuan Lyu, Hanyang Liu, Jianbo Song et al.AAAI 2026
- Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object DetectionWenjun Hui, Zhenfeng Zhu, Shuai Zheng, Yao ZhaoCVPR 2024
- SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object TrackingYangkai Chen, Qiangqiang Wu, Guangyao Li, Junlong Gao et al.AAAI 2026
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
