Learning to Compose: Improving Object Centric Learning by Injecting Compositionality
Whie Jung, Jaehoon Yoo, Sungjin Ahn, Seunghoon Hong
Abstract
Learning compositional representation is a key aspect of object-centric learning as it enables flexible systematic generalization and supports complex visual reasoning. However, most of the existing approaches rely on auto-encoding objective, while the compositionality is implicitly imposed by the architectural or algorithmic bias in the encoder. This misalignment between auto-encoding objective and learning compositionality often results in failure of capturing meaningful object representations. In this study, we propose a novel objective that explicitly encourages compositionality of the representations. Built upon the existing object-centric learning framework (e.g., slot attention), our method incorporates additional constraints that an arbitrary mixture of object representations from two images should be valid by maximizing the likelihood of the composite data. We demonstrate that incorporating our objective to the existing framework consistently improves the objective-centric learning and enhances the robustness to the architectural choices. Codes are available at https://github.com/whieya/Learning-to-compose .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?Yihao Li, Saeed Salehi, Lyle H. Ungar, Konrad P. KordingNeurIPS 2025 · 20 citations
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai et al.ICLR 2026 · 3 citations
- Disentangled Representation Learning via Modular Compositional BiasWhie Jung, Dong Hoon Lee, Seunghoon HongNeurIPS 2025 · 3 citations
- Is Generation Required for Data-Efficient Perception?Jack Brady, Bernhard Schölkopf, Thomas Kipf, Simon Buchholz et al.ICML 2026 · 1 citation
- Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence AnalysisBoyang Dai, Chaoqi Chen, Yizhou YuCVPR 2026 · 1 citation
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
Related papers
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 29 citations
- Provable Compositional Generalization for Object-Centric LearningThaddäus Wiedemer, Jack Brady, Alexander Panfilov, Attila Juhos et al.ICLR 2024 · 40 citations
- Identifiable Object-Centric Representation Learning via Probabilistic Slot AttentionAvinash Kori, Francesco Locatello, Ainkaran Santhirasekaram, Francesca Toni et al.NeurIPS 2024 · 11 citations
- Adaptive Slot Attention: Object Discovery with Dynamic Slot NumberKe Fan, Zechen Bai, Tianjun Xiao, Tong He et al.CVPR 2024
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
