Bootstrapping Objectness from Videos by Relaxed Common Fate and Visual Grouping
Long Lian, Zhirong Wu, Stella X. Yu
摘要
We study learning object segmentation from unlabeled videos. Humans can easily segment moving objects without knowing what they are. The Gestalt law of common fate, i.e., what move at the same speed belong together, has inspired unsupervised object discovery based on motion segmentation. However, common fate is not a reliable indicator of objectness: Parts of an articulated / deformable object may not move at the same speed, whereas shadows / reflections of an object always move with it but are not part of it. Our insight is to bootstrap objectness by first learning image features from relaxed common fate and then refining them based on visual appearance grouping within the image itself and across images statistically. Specifically, we learn an image segmenter first in the loop of approximating optical flow with constant segment flow plus small withinsegment residual flow, and then by refining it for more coherent appearance and statistical figure-ground relevance. On unsupervised video object segmentation, using only ResNet and convolutional heads, our model surpasses the state-of-the-art by absolute gains of 7/9/5% on DAVIS16 / STv2 / FBMS59 respectively, demonstrating the effectiveness of our ideas. Our code is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Learning Segmentation from Point TrajectoriesLaurynas Karazija, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2024 · 被引用 14 次
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICCV 2025 · 被引用 11 次
- GeoMotion: Rethinking Motion Segmentation via Latent 4D GeometryXiankang He, Peile Lin, Ying Cui, Dongyan Guo 等CVPR 2026 · 被引用 2 次
- SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMsChong Tang, Sannara Ek, Dirk Koch, Robert Mullins 等ICLR 2026
- Segment Any Motion in VideosNan Huang, Wenzhao Zheng, Chenfeng Xu, Kurt Keutzer 等CVPR 2025
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Rethinking Space-Time Networks with Improved Memory Coverage for Efficient Video Object SegmentationHo Kei Cheng, Yu-Wing Tai, Chi-Keung TangNeurIPS 2021 · 被引用 403 次
相关 Paper
- The Emergence of Objectness: Learning Zero-shot Segmentation from VideosRuntao Liu, Zhirong Wu, Stella X. Yu, Stephen LinNeurIPS 2021 · 被引用 62 次
- DyStaB: Unsupervised Object Segmentation via Dynamic-Static BootstrappingYanchao Yang, Brian Lai, Stefano SoattoCVPR 2021
- SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationLingyi Hong, Wei Zhang, Shuyong Gao, Hong Lu 等ACM MM 2023 · 被引用 14 次
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman 等ICCV 2021 · 被引用 188 次
- Target-Aware Object Discovery and Association for Unsupervised Video Multi-Object SegmentationTianfei Zhou, Jianwu Li, Xueyi Li, Ling ShaoCVPR 2021
