Pay Attention to the Foreground in Object-Centric Learning
Pinzhuo Tian, Shengjie Yang, Hang Yu, Alex C. Kot
Abstract
The slot attention-based method is widely used in unsupervised object-centric learning, aiming to decompose scenes into interpretable objects and associate them with slots. However, complex backgrounds in the real images can disrupt the model's focus, leading it to excessively segment background stuff into different regions based on lowlevel information such as color or texture variations. As a result, the detailed segmentation of foreground objects, which requires shape or geometric information, is often neglected. To address this issue, we introduce a contrastive learning-based indicator designed to differentiate between foreground and background. Integrating this indicator into the existing slot attention-based method enables the model to focus more on segmenting foreground objects while minimizing background distractions. During the testing phase, we utilize a spectral clustering mechanism to refine the results based on the similarity between the slots. Experimental results show that incorporating our method with various state-of-the-art models significantly improves their performance on both simulated data and real-world datasets. Furthermore, multiple sets of ablation experiments confirm the effectiveness of each proposed component. The source code is available at https://github.com/ sjyjs09/FG-BG_Indicator .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3554504-56e5-4fde-acbe-c3329dfb05cdCited by top-tier papers1
Ask how each one uses itBuilds on22
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Dense Visual RepresentationsPedro O. Pinheiro, Amjad Almahairi, Ryan Y. Benmalek, Florian Golemo et al.NeurIPS 2020 · 227 citations
- Kubric: A scalable dataset generatorKlaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch et al.CVPR 2022 · 183 citations
Related papers
- Self-Supervised Visual Representation Learning with Semantic GroupingXin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang et al.NeurIPS 2022 · 104 citations
- Guided Slot Attention for Unsupervised Video Object SegmentationMinhyeok Lee, Suhwan Cho, Dogyoon Lee, Chaewon Park et al.CVPR 2024
- Distilling Localization for Self-Supervised Representation LearningNanxuan Zhao, Zhirong Wu, Rynson W. H. Lau, Stephen LinAAAI 2021 · 59 citations
- FOCUS: Towards Universal Foreground SegmentationZuyao You, Lingyu Kong, Lingchen Meng, Zuxuan WuAAAI 2025 · 9 citations
- Improving Object-centric Learning with Query OptimizationBaoxiong Jia, Yu Liu, Siyuan HuangICLR 2023 · 4 citations
