Multi-Scale Fusion for Object Representation
Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen
Abstract
Representing images or videos as object-level feature vectors, rather than pixellevel feature maps, facilitates advanced visual tasks. Object-Centric Learning (OCL) primarily achieves this by reconstructing the input under the guidance of Variational Autoencoder (VAE) intermediate representation to drive so-called slots to aggregate as much object information as possible. However, existing VAE guidance does not explicitly address that objects can vary in pixel sizes while models typically excel at specific pattern scales. We propose Multi-Scale Fusion (MSF) to enhance VAE guidance for OCL training. To ensure objects of all sizes fall within VAE's comfort zone, we adopt the image pyramid, which produces intermediate representations at multiple scales; To foster scale-invariance/variance in object super-pixels, we devise inter/intra-scale fusion, which augments lowquality object super-pixels of one scale with corresponding high-quality superpixels from another scale. On standard OCL benchmarks, our technique improves mainstream methods, including state-of-the-art diffusion-based ones. The source code is available on https://github.com/Genera1Z/MultiScaleFusion .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Smoothing Slot Attention Iterations and RecurrencesRongzhen Zhao, Wenyan Yang, Kannala Juho, Joni PajarinenICML 2026 · 4 citations
- Predicting Video Slot Attention Queries from Random Slot-Feature PairsRongzhen Zhao, Jian Li, Juho Kannala, Joni PajarinenAAAI 2026 · 3 citations
- Slot Attention with Re-Initialization and Self-DistillationRongzhen Zhao, Yi Zhao, Juho Kannala, Joni PajarinenACM MM 2025 · 1 citation
- Vector-Quantized Vision Foundation Models for Object-Centric LearningRongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni PajarinenACM MM 2025
Builds on21
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks With Octave ConvolutionYunpeng Chen, Haoqi Fan, Bing Xu, Zhicheng Yan et al.ICCV 2019 · 665 citations
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
Related papers
- MetaSlot: Break Through the Fixed Number of Slots in Object-Centric LearningHongjia Liu, Rongzhen Zhao, Haohan Chen, Joni PajarinenNeurIPS 2025 · 12 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- MUFASA: A Multi-Layer Framework for Slot AttentionSebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer et al.CVPR 2026 · 1 citation
- Self-supervised Object-Centric Learning for VideosGörkay Aydemir, Weidi Xie, Fatma GüneyNeurIPS 2023 · 61 citations
- From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric LearningHyun Seok Seong, WonJun Moon, Jae-Pil HeoICLR 2026 · 5 citations
