Robust and Controllable Object-Centric Learning through Energy-based Models
Ruixiang Zhang, Tong Che, Boris Ivanovic, Renhao Wang, Marco Pavone, Yoshua Bengio, Liam Paull
Abstract
Humans are remarkably good at understanding and reasoning about complex visual scenes. The capability to decompose low-level observations into discrete objects allows us to build a grounded abstract representation and identify the compositional structure of the world. Accordingly, it is a crucial step for machine learning models to be capable of inferring objects and their properties from visual scenes without explicit supervision. However, existing works on objectcentric representation learning either rely on tailor-made neural network modules or strong probabilistic assumptions in the underlying generative and inference processes. In this work, we present EGO, a conceptually simple and general approach to learning object-centric representations through an energy-based model. By forming a permutation-invariant energy function using vanilla attention blocks readily available in Transformers, we can infer object-centric latent variables via gradient-based MCMC methods where permutation equivariance is automatically guaranteed. We show that EGO can be easily integrated into existing architectures and can effectively extract high-quality object-centric representations, leading to better segmentation accuracy and competitive downstream task performance. Further, empirical evaluations show that EGO's learned representations are robust against distribution shift. Finally, we demonstrate the effectiveness of EGO in systematic compositional generalization, by re-composing learned energy functions for novel scene generation and manipulation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a20d0eb-5c69-437f-a3c7-7e8194e8b76fCited by top-tier papers6
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 106 citations
- Rotating Features for Object DiscoverySindy Löwe, Phillip Lippe, Francesco Locatello, Max WellingNeurIPS 2023 · 37 citations
- Neural Language of Thought ModelsYi-Fu Wu, Minseung Lee, Sungjin AhnICLR 2024 · 11 citations
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang et al.NeurIPS 2025 · 8 citations
- Inferring Relational Potentials in Interacting SystemsArmand Comas Massague, Yilun Du, Christian Fernandez Lopez, Sandesh Ghimire et al.ICML 2023 · 6 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 334 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
Related papers
- Provably Learning Object-Centric RepresentationsJack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf et al.ICML 2023 · 55 citations
- Generalization and Robustness Implications in Object-Centric LearningAndrea Dittadi, Samuele S. Papa, Michele De Vita, Bernhard Schölkopf et al.ICML 2022 · 87 citations
- Bridging the Gap to Real-World Object-Centric LearningMaximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow et al.ICLR 2023 · 31 citations
- Unsupervised Causal Generative Understanding of ImagesTitas Anciukevicius, Patrick Fox-Roberts, Edward Rosten, Paul HendersonNeurIPS 2022 · 6 citations
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun et al.NeurIPS 2021 · 17 citations
