Self-supervised Video Object Segmentation by Motion Grouping
Charig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman, Weidi Xie
摘要
Animals have evolved highly functional visual systems to understand motion, assisting perception even under complex environments. In this paper, we work towards developing a computer vision system able to segment objects by exploiting motion cues, i.e. motion segmentation. To achieve this, we introduce a simple variant of the Transformer to segment optical flow frames into primary objects and the background, which can be trained in a self-supervised manner, i.e. without using any manual annotations. Despite using only optical flow, and no appearance information, as input, our approach achieves superior results compared to previous state-of-the-art self-supervised methods on public benchmarks (DAVIS2016, SegTrackv2, FBMS59), while being an order of magnitude faster. On a challenging camouflage dataset (MoCA), we significantly outperform other self-supervised approaches, and are competitive with the top supervised approach, highlighting the importance of motion cues and the potential bias towards appearance in existing video segmentation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper60
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- SAVi++: Towards End-to-End Object-Centric Learning from Real-World VideosGamaleldin F. Elsayed, Aravindh Mahendran, Sjoerd van Steenkiste, Klaus Greff 等NeurIPS 2022 · 被引用 218 次
- D^2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular VideoTianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, Forrester Cole 等NeurIPS 2022 · 被引用 184 次
- CRAFT: Cross-Attentional Flow Transformer for Robust Optical FlowXiuchao Sui, Shaohua Li, Xue Geng, Yan Wu 等CVPR 2022 · 被引用 114 次
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski 等NeurIPS 2023 · 被引用 106 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 被引用 74 次
- Implicit Motion Handling for Video Camouflaged Object DetectionXuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong 等CVPR 2022 · 被引用 83 次
- Video Diffusion Models Excel at Tracking Similar-Looking Objects Without SupervisionChenshuang Zhang, Kang Zhang, Joon Son Chung, In So Kweon 等NeurIPS 2025
- Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual GroupingLong Lian, Zhirong Wu, Stella X. YuCVPR 2023
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu 等ACM MM 2023 · 被引用 14 次
