Self-supervised Video Object Segmentation by Motion Grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, Weidi Xie
Abstract
Animals have evolved highly functional visual systems to understand motion, assisting perception even under complex environments. In this paper, we work towards developing a computer vision system able to segment objects by exploiting motion cues, i.e. motion segmentation. To achieve this, we introduce a simple variant of the Transformer to segment optical flow frames into primary objects and the background, which can be trained in a self-supervised manner, i.e. without using any manual annotations. Despite using only optical flow, and no appearance information, as input, our approach achieves superior results compared to previous state-of-the-art self-supervised methods on public benchmarks (DAVIS2016, SegTrackv2, FBMS59), while being an order of magnitude faster. On a challenging camouflage dataset (MoCA), we significantly outperform other self-supervised approaches, and are competitive with the top supervised approach, highlighting the importance of motion cues and the potential bias towards appearance in existing video segmentation models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc264a24-cf2b-4bdd-8e8f-19b2f0ebc1f9Cited by top-tier papers60
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
- SAVi++: Towards End-to-End Object-Centric Learning from Real-World VideosGamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff et al.NeurIPS 2022 · 218 citations
- D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular VideoTianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, Forrester Cole et al.NeurIPS 2022 · 184 citations
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowXiuchao Sui, Shaohua Li, Xue Geng, Yan Wu et al.CVPR 2022 · 114 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
Related papers
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 74 citations
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong et al.CVPR 2022 · 83 citations
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionChenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon et al.NeurIPS 2025
- Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingLong Lian, Zhirong Wu, Stella X. YuCVPR 2023
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
