Slot-Guided Adaptation of Pre-trained Diffusion Models for Object-Centric Learning and Compositional Generation
Adil Kaan Akan, Yucel Yemez
Abstract
We present SlotAdapt, an object-centric learning method that combines slot attention with pretrained diffusion models by introducing adapters for slot-based conditioning. Our method preserves the generative power of pretrained diffusion models, while avoiding their text-centric conditioning bias. We also incorporate an additional guidance loss into our architecture to align cross-attention from adapter layers with slot attention. This enhances the alignment of our model with the objects in the input image without using external supervision. Experimental results show that our method outperforms state-of-the-art techniques in object discovery and image generation tasks across multiple datasets, including those with real images. Furthermore, we demonstrate through experiments that our method performs remarkably well on complex real-world images for compositional generation, in contrast to other slot-based generative methods in the literature. The project page can be found at https://kaanakan.github.io/SlotAdapt/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2f6b29a-1a9d-4fd0-a07d-b5c00dd046e4Cited by top-tier papers6
- Object-Centric Latent Action LearningAlbina Klepach, Alexander Nikulin, Ilya Zisman, Denis Tarasov et al.AAAI 2026 · 7 citations
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai et al.ICLR 2026 · 3 citations
- MUFASA: A Multi-Layer Framework for Slot AttentionSebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer et al.CVPR 2026 · 1 citation
- A Compressive-Expressive Communication Framework for Compositional RepresentationsRafael Elberg, Felipe del Río, Mircea Petrache, Denis ParraNeurIPS 2025 · 1 citation
- Factor-Wise Homogeneity of Slot-Attention for Continual Object-Centric LearningIlmin Kang, Hoyong Kim, Seungju Bang, Minwoo Kang et al.ICML 2026
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- GLASS: Guided Latent Slot Diffusion for Object-Centric LearningKrishnakant Singh, Simone Schaub-Meyer, Stefan RothCVPR 2025
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 106 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- Slot-VAE: Object-Centric Scene Generation with Slot AttentionYanbo Wang, Letao Liu, Justin DauwelsICML 2023 · 29 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
