Shepherding Slots to Objects: Towards Stable and Robust Object-Centric Learning
Jinwoo Kim, Janghyuk Choi, Ho-Jin Choi, Seon Joo Kim
摘要
Object-centric learning (OCL) aspires general and compositional understanding of scenes by representing a scene as a collection of object-centric representations. OCL has also been extended to multi-view image and video datasets to apply various data-driven inductive biases by utilizing geometric or temporal information in the multi-image data. Single-view images carry less information about how to disentangle a given scene than videos or multi-view images do. Hence, owing to the difficulty of applying inductive biases, OCL for single-view images remains challenging, resulting in inconsistent learning of object-centric representation. To this end, we introduce a novel OCL framework for single-view images, SLot Attention via SHepherding (SLASH), which consists of two simple-yet-effective modules on top of Slot Attention. The new modules, Attention Refining Kernel (ARK) and Intermediate Point Predictor and Encoder (IPPE), respectively, prevent slots from being distracted by the background noise and indicate locations for slots to focus on to facilitate learning of objectcentric representation. We also propose a weak semisupervision approach for OCL, whilst our proposed framework can be used without any assistant annotation during the inference. Experiments show that our proposed method enables consistent learning of object-centric representation and achieves strong performance across four datasets. Code is available at https://github.com/object- understanding/SLASH.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 被引用 104 次
- Learning to Compose: Improving Object Centric Learning by Injecting CompositionalityWhie Jung, Jaehoon Yoo, Sungjin Ahn, Seunghoon HongICLR 2024 · 被引用 10 次
- Bootstrapping Top-down Information for Self-modulating Slot AttentionDongwon Kim, Seoyeon Kim, Suha KwakNeurIPS 2024 · 被引用 7 次
- Multi-Object Representation Learning via Feature Connectivity and Object-Centric RegularizationAlex Foo, Wynne Hsu, Mong-Li LeeNeurIPS 2023 · 被引用 4 次
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper19
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- End-to-End Semi-Supervised Object Detection with Soft TeacherMengde Xu, Zheng Zhang, Han Hu, Jianfeng Wang 等ICCV 2021 · 被引用 622 次
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 被引用 334 次
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- SPACE: Unsupervised Object-Oriented Scene Representation via Spatial Attention and DecompositionZhixuan Lin, Yi-Fu Wu, Skand Vishwanath Peri, Weihao Sun 等ICLR 2020 · 被引用 276 次
相关 Paper
- Slot Attention with Re-Initialization and Self-DistillationRongzhen Zhao, Yi Zhao, Juho Kannala, Joni PajarinenACM MM 2025 · 被引用 1 次
- Smoothing Slot Attention Iterations and RecurrencesRongzhen Zhao, Wenyan Yang, Kannala Juho, Joni PajarinenICML 2026 · 被引用 4 次
- Simple Unsupervised Object-Centric Learning for Complex and Naturalistic VideosGautam Singh, Yi-Fu Wu, Sungjin AhnNeurIPS 2022 · 被引用 182 次
- Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric LearningWonJun Moon, Hyun Seok Seong, Jae-Pil HeoCVPR 2026 · 被引用 3 次
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 被引用 29 次
