Robust and Controllable Object-Centric Learning through Energy-based Models
Ruixiang Zhang, Tong Che, Boris Ivanovic, Renhao Wang, Marco Pavone, Yoshua Bengio, Liam Paull
摘要
Humans are remarkably good at understanding and reasoning about complex visual scenes. The capability to decompose low-level observations into discrete objects allows us to build a grounded abstract representation and identify the compositional structure of the world. Accordingly, it is a crucial step for machine learning models to be capable of inferring objects and their properties from visual scenes without explicit supervision. However, existing works on objectcentric representation learning either rely on tailor-made neural network modules or strong probabilistic assumptions in the underlying generative and inference processes. In this work, we present EGO, a conceptually simple and general approach to learning object-centric representations through an energy-based model. By forming a permutation-invariant energy function using vanilla attention blocks readily available in Transformers, we can infer object-centric latent variables via gradient-based MCMC methods where permutation equivariance is automatically guaranteed. We show that EGO can be easily integrated into existing architectures and can effectively extract high-quality object-centric representations, leading to better segmentation accuracy and competitive downstream task performance. Further, empirical evaluations show that EGO's learned representations are robust against distribution shift. Finally, we demonstrate the effectiveness of EGO in systematic compositional generalization, by re-composing learned energy functions for novel scene generation and manipulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 被引用 106 次
- Rotating Features for Object DiscoverySindy Löwe, Phillip Lippe, Francesco Locatello, Max WellingNeurIPS 2023 · 被引用 37 次
- Neural Language of Thought ModelsYi-Fu Wu, Minseung Lee, Sungjin AhnICLR 2024 · 被引用 11 次
- Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive FlowsRuixiang Zhang, Shuangfei Zhai, Jiatao Gu, Yizhe Zhang 等NeurIPS 2025 · 被引用 8 次
- Inferring Relational Potentials in Interacting SystemsArmand Comas Massague, Yilun Du, Christian Fernandez Lopez, Sandesh Ghimire 等ICML 2023 · 被引用 6 次
它引用的顶会 Paper16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent RepresentationsMartin Engelcke, Adam R. Kosiorek, Oiwi Parker Jones, Ingmar PosnerICLR 2020 · 被引用 334 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
相关 Paper
- Provably Learning Object-Centric RepresentationsJack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf 等ICML 2023 · 被引用 55 次
- Generalization and Robustness Implications in Object-Centric LearningAndrea Dittadi, Samuele S. Papa, Michele De Vita, Bernhard Schölkopf 等ICML 2022 · 被引用 87 次
- Bridging the Gap to Real-World Object-Centric LearningMaximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow 等ICLR 2023 · 被引用 31 次
- Unsupervised Causal Generative Understanding of ImagesTitas Anciukevicius, Patrick Fox-Roberts, Edward Rosten, Paul HendersonNeurIPS 2022 · 被引用 6 次
- Object-Centric Representation Learning with Generative Spatial-Temporal FactorizationNanbo Li, Muhammad Ahmed Raza, Wenbin Hu, Zhaole Sun 等NeurIPS 2021 · 被引用 17 次
