MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
Hongjia Liu, Rongzhen Zhao, Haohan Chen, Joni Pajarinen
摘要
Learning object-level, structured representations is widely regarded as a key to better generalization in vision and underpins the design of next-generation Pre-trained Vision Models (PVMs). Mainstream Object-Centric Learning (OCL) methods adopt Slot Attention or its variants to iteratively aggregate objects' super-pixels into a fixed set of query feature vectors, termed slots. However, their reliance on a static slot count leads to an object being represented as multiple parts when the number of objects varies. We introduce MetaSlot, a plug-and-play Slot Attention variant that adapts to variable object counts. MetaSlot (i) maintains a codebook that holds prototypes of objects in a dataset by vector-quantizing the resulting slot representations; (ii) removes duplicate slots from the traditionally aggregated slots by quantizing them with the codebook; and (iii) injects progressively weaker noise into the Slot Attention iterations to accelerate and stabilize the aggregation. MetaSlot is a general Slot Attention variant that can be seamlessly integrated into existing OCL architectures. Across multiple public datasets and tasks-including object discovery and recognition-models equipped with MetaSlot achieve significant performance gains and markedly interpretable slot representations, compared with existing Slot Attention variants. The code is available at https://github.com/lhj-lhj/MetaSlot.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric LearningHyun Seok Seong, WonJun Moon, Jae-Pil HeoICLR 2026 · 被引用 5 次
- Smoothing Slot Attention Iterations and RecurrencesRongzhen Zhao, Wenyan Yang, Kannala Juho, Joni PajarinenICML 2026 · 被引用 4 次
- Factor-Wise Homogeneity of Slot-Attention for Continual Object-Centric LearningIlmin Kang, Hoyong Kim, Seungju Bang, Minwoo Kang 等ICML 2026
- Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to CorrespondenceZhiyuan Li, Rongzhen Zhao, Wenyan Yang, Wenshuai Zhao 等ICML 2026
它引用的顶会 Paper40
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang 等ICLR 2022 · 被引用 753 次
相关 Paper
- Slot Attention with Re-Initialization and Self-DistillationRongzhen Zhao, Yi Zhao, Juho Kannala, Joni PajarinenACM MM 2025 · 被引用 1 次
- Adaptive Slot Attention: Object Discovery with Dynamic Slot NumberKe Fan, Zechen Bai, Tianjun Xiao, Tong He 等CVPR 2024
- MUFASA: A Multi-Layer Framework for Slot AttentionSebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer 等CVPR 2026 · 被引用 1 次
- Vector-Quantized Vision Foundation Models for Object-Centric LearningRongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni PajarinenACM MM 2025
- Grounded Object-Centric LearningAvinash Kori, Francesco Locatello, Fabio De Sousa Ribeiro, Francesca Toni 等ICLR 2024 · 被引用 17 次
