Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
Rui Qian, Shuangrui Ding, Xian Liu, Dahua Lin
摘要
Self-supervised methods have shown remarkable progress in learning high-level semantics and low-level temporal correspondence. Building on these results, we take one step further and explore the possibility of integrating these two features to enhance object-centric representations. Our preliminary experiments indicate that query slot attention can extract different semantic components from the RGB feature map, while random sampling based slot attention can exploit temporal correspondence cues between frames to assist instance identification. Motivated by this, we propose a novel semantic-aware masked slot attention on top of the fused semantic features and correspondence maps. It comprises two slot attention stages with a set of shared learnable Gaussian distributions. In the first stage, we use the mean vectors as slot initialization to decompose potential semantics and generate semantic segmentation masks through iterative attention. In the second stage, for each semantics, we randomly sample slots from the corresponding Gaussian distribution and perform masked feature aggregation within the semantic area to exploit temporal correspondence patterns for instance identification. We adopt semantic- and instance-level temporal consistency as self-supervision to encourage temporally coherent object-centric representations. Our model effectively identifies multiple object instances with semantic structure, reaching promising results on unsupervised video object discovery. Furthermore, we achieve state-of-the-art performance on dense label propagation tasks, demonstrating the potential for object-centric analysis. The code is released at https://github.com/shvdiwnkozbw/SMTC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 被引用 104 次
- Emergent Temporal Correspondences from Video Diffusion TransformersJisu Nam, Soowon Son, Dahyun Chung, Jiyoung Kim 等NeurIPS 2025 · 被引用 30 次
- Advancing Complex Video Object Segmentation via Progressive Concept ConstructionZhixiong Zhang, Shuangrui Ding, Xiaoyi Dong, Songxin He 等ICLR 2026 · 被引用 17 次
- SAM2LONG: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory TreeShuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang 等ICCV 2025 · 被引用 15 次
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi 等ICLR 2024 · 被引用 11 次
它引用的顶会 Paper36
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
相关 Paper
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius 等CVPR 2025
- Guided Slot Attention for Unsupervised Video Object SegmentationMinhyeok Lee, Suhwan Cho, Dogyoon Lee, Chaewon Park 等CVPR 2024
- Unsupervised Open-Vocabulary Object Localization in VideosKe Fan, Zechen Bai, Tianjun Xiao, Dominik Zietlow 等ICCV 2023 · 被引用 14 次
- Smoothing Slot Attention Iterations and RecurrencesRongzhen Zhao, Wenyan Yang, Kannala Juho, Joni PajarinenICML 2026 · 被引用 4 次
