Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
Amir Mohammad Karimi-Mamaghan, Samuele Papa, Karl Henrik Johansson, Stefan Bauer, Andrea Dittadi
摘要
Object-centric (OC) representations, which model visual scenes as compositions of discrete objects, have the potential to be used in various downstream tasks to achieve systematic compositional generalization and facilitate reasoning. However, these claims have yet to be thoroughly validated empirically. Recently, foundation models have demonstrated unparalleled capabilities across diverse domains, from language to computer vision, positioning them as a potential cornerstone of future research for a wide range of computational tasks. In this paper, we conduct an extensive empirical study on representation learning for downstream Visual Question Answering (VQA), which requires an accurate compositional understanding of the scene. We thoroughly investigate the benefits and trade-offs of OC models and alternative approaches including large pre-trained foundation models on both synthetic and real-world data, ultimately identifying a promising path to leverage the strengths of both paradigms. The extensiveness of our study, encompassing over 600 downstream VQA models and 15 different types of upstream representations, also provides several additional insights that we believe will be of interest to the community at large.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Object-Centric Concept-BottlenecksDavid Steinmann, Wolfgang Stammer, Antonia Wüst, Kristian KerstingNeurIPS 2025 · 被引用 12 次
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual RecallJiayu Yang, Yuxuan Fan, Songning Lai, Shengen Wu 等ICLR 2026 · 被引用 5 次
- CTRL-O: Language-Controllable Object-Centric Visual Representation LearningAniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer 等CVPR 2025
- Temporally Consistent Object-Centric Learning by Contrasting SlotsAnna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius 等CVPR 2025
- PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and PlanningAngel Villar-Corrales, Sven BehnkeICML 2025
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Are Object-Centric Representations Better at Compositional Generalization?Ferdinand Kapl, Amir Mohammad Karimi Mamaghan, Maximilian Seitzer, Karl Johansson 等ICML 2026
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris 等CVPR 2021
- Evaluating Object-Centric Models beyond Object DiscoveryKrishnakant Singh, Simone Schaub-Meyer, Stefan RothICML 2026
- Vector-Quantized Vision Foundation Models for Object-Centric LearningRongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni PajarinenACM MM 2025
- Probing the 3D Awareness of Visual Foundation ModelsMohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar 等CVPR 2024
