SlotDiffusion: Object-Centric Generative Modeling with Diffusion Models
Ziyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski, Animesh Garg
摘要
Object-centric learning aims to represent visual data with a set of object entities (a.k.a. slots), providing structured representations that enable systematic generalization. Leveraging advanced architectures like Transformers, recent approaches have made significant progress in unsupervised object discovery. In addition, slot-based representations hold great potential for generative modeling, such as controllable image generation and object manipulation in image editing. However, current slot-based methods often produce blurry images and distorted objects, exhibiting poor generative modeling capabilities. In this paper, we focus on improving slot-toimage decoding, a crucial aspect for high-quality visual generation. We introduce SlotDiffusion -an object-centric Latent Diffusion Model (LDM) designed for both image and video data. Thanks to the powerful modeling capacity of LDMs, SlotDiffusion surpasses previous slot models in unsupervised object segmentation and visual generation across six datasets. Furthermore, our learned object features can be utilized by existing object-centric dynamics models, improving video prediction quality and downstream temporal reasoning tasks. Finally, we demonstrate the scalability of SlotDiffusion to unconstrained real-world datasets such as PASCAL VOC and COCO, when integrated with self-supervised pre-trained image encoders. Additional results and details are available at our website.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper43
- Object-Centric Slot DiffusionJindong Jiang, Fei Deng, Gautam Singh, Sungjin AhnNeurIPS 2023 · 被引用 106 次
- Object-Centric Learning for Real-World Videos by Predicting Temporal Feature SimilaritiesAndrii Zadaianchuk, Maximilian Seitzer, Georg MartiusNeurIPS 2023 · 被引用 104 次
- Diffusion Model with Cross Attention as an Inductive Bias for DisentanglementTao Yang, Cuiling Lan, Yan Lu, Nanning ZhengNeurIPS 2024 · 被引用 41 次
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson 等NeurIPS 2024 · 被引用 32 次
- Slot State Space ModelsJindong Jiang, Fei Deng, Gautam Singh, Minseung Lee 等NeurIPS 2024 · 被引用 18 次
它引用的顶会 Paper64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- GLASS: Guided Latent Slot Diffusion for Object-Centric LearningKrishnakant Singh, Simone Schaub-Meyer, Stefan RothCVPR 2025
- Slot-Guided Adaptation of Pre-trained Diffusion Models for Object-Centric Learning and Compositional GenerationAdil Kaan Akan, Yucel YemezICLR 2025
- Improved Object-Centric Diffusion Learning with Registers and Contrastive AlignmentBac Nguyen, Yuhta Takida, Naoki Murata, Chieh-Hsin Lai 等ICLR 2026 · 被引用 3 次
- SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric ModelsZiyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf 等ICLR 2023 · 被引用 10 次
- CTRL-O: Language-Controllable Object-Centric Visual Representation LearningAniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer 等CVPR 2025
