Self-supervised Object-Centric Learning for Videos
Görkay Aydemir, Weidi Xie, Fatma Güney
摘要
Unsupervised multi-object segmentation has shown impressive results on images by utilizing powerful semantics learned from self-supervised pretraining. An additional modality such as depth or motion is often used to facilitate the segmentation in video sequences. However, the performance improvements observed in synthetic sequences, which rely on the robustness of an additional cue, do not translate to more challenging real-world scenarios. In this paper, we propose the first fully unsupervised method for segmenting multiple objects in real-world sequences. Our object-centric learning framework spatially binds objects to slots on each frame and then relates these slots across frames. From these temporally-aware slots, the training objective is to reconstruct the middle frame in a high-level semantic feature space. We propose a masking strategy by dropping a significant portion of tokens in the feature space for efficiency and regularization. Additionally, we address over-clustering by merging slots based on similarity. Our method can successfully segment multiple instances of complex and high-variety classes in YouTube videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 被引用 104 次
- Dyn-O: Building Structured World Models with Object-Centric RepresentationsZizhao Wang, Kaixin Wang, Li Zhao, Peter Stone 等NeurIPS 2025 · 被引用 15 次
- Learning Segmentation from Point TrajectoriesLaurynas Karazija, Iro Laina, Christian Rupprecht, Andrea VedaldiNeurIPS 2024 · 被引用 14 次
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric LearningHongjia Liu, Rongzhen Zhao, Haohan Chen, Joni PajarinenNeurIPS 2025 · 被引用 12 次
- Causal-JEPA: Learning World Models through Object-Level Latent MaskingHeejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun 等ICML 2026 · 被引用 7 次
它引用的顶会 Paper39
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 被引用 615 次
相关 Paper
- Object-Centric Multiple Object TrackingZixu Zhao, Jiaze Wang, Max Horn, Yizhuo Ding 等ICCV 2023 · 被引用 10 次
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius 等CVPR 2025
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 被引用 182 次
- Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in VideosRui Qian, Shuangrui Ding, Xian Liu, Dahua LinICCV 2023 · 被引用 23 次
