Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual Grouping
Long Lian, Zhirong Wu, Stella X. Yu
Abstract
We study learning object segmentation from unlabeled videos. Humans can easily segment moving objects without knowing what they are. The Gestalt law of common fate, i.e., what move at the same speed belong together, has inspired unsupervised object discovery based on motion segmentation. However, common fate is not a reliable indicator of objectness: Parts of an articulated / deformable object may not move at the same speed, whereas shadows / reflections of an object always move with it but are not part of it. Our insight is to bootstrap objectness by first learning image features from relaxed common fate and then refining them based on visual appearance grouping within the image itself and across images statistically. Specifically, we learn an image segmenter first in the loop of approximating optical flow with constant segment flow plus small withinsegment residual flow, and then by refining it for more coherent appearance and statistical figure-ground relevance. On unsupervised video object segmentation, using only ResNet and convolutional heads, our model surpasses the state-of-the-art by absolute gains of 7/9/5% on DAVIS16 / STv2 / FBMS59 respectively, demonstrating the effectiveness of our ideas. Our code is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d98bfdc5-c0b8-4cb3-8f3a-72b6ca6c06d9Cited by top-tier papers6
- Learning Segmentation from Point TrajectoriesLaurynas Karazija, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2024 · 14 citations
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger et al.ICCV 2025 · 11 citations
- GeoMotion: Rethinking Motion Segmentation via Latent 4D GeometryXiankang He, Peile Lin, Ying Cui, Dongyan Guo et al.CVPR 2026 · 2 citations
- SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMsChong Tang, Sannara Ek, Dirk Koch, Robert Mullins et al.ICLR 2026
- Segment Any Motion in VideosNan Huang, Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer et al.CVPR 2025
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 403 citations
Related papers
- The Emergence of Objectness: Learning Zero-shot Segmentation from VideosRuntao Liu, Zhirong Wu, Stella X. Yu, Stephen LinNeurIPS 2021 · 62 citations
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu et al.ACM MM 2023 · 14 citations
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object SegmentationTianfei Zhou, Jianwu Li, Xueyi Li, Ling ShaoCVPR 2021
