Attention over Learned Object Embeddings Enables Complex Visual Reasoning
David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, Matt M. Botvinick
Abstract
Neural networks have achieved success in a wide array of perceptual tasks but often fail at tasks involving both perception and higher-level reasoning. On these more challenging tasks, bespoke approaches (such as modular symbolic components, independent dynamics models or semantic parsers) targeted towards that specific type of task have typically performed better. The downside to these targeted approaches, however, is that they can be more brittle than general-purpose neural networks, requiring significant modification or even redesign according to the particular task at hand. Here, we propose a more general neural-network-based approach to dynamic visual reasoning problems that obtains state-of-the-art performance on three different domains, in each case outperforming bespoke modular approaches tailored specifically to the task. Our method relies on learned object-centric representations, self-attention and self-supervised dynamics learning, and all three elements together are required for strong performance to emerge. The success of this combination suggests that there may be no need to trade off flexibility for performance on problems involving spatio-temporal or causal-style reasoning. With the right soft biases and learning objectives in a neural network we may be able to attain the best of both worlds.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53333389-00e8-4e2c-8e28-8a5c5bd554f8Cited by top-tier papers29
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
- SlotDiffusion: Object-Centric Generative Modeling with Diffusion ModelsZiyi Wu, Jingyu Hu, Wuyue Lu, Igor Gilitschenski et al.NeurIPS 2023 · 106 citations
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata et al.NeurIPS 2024 · 101 citations
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain et al.NeurIPS 2025 · 90 citations
- OGC: Unsupervised 3D Object Segmentation from Rigid Dynamics of Point CloudsZiyang Song, Bo YangNeurIPS 2022 · 41 citations
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
Related papers
- NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-LanguageDanial Kamali, Parisa KordjamshidiICLR 2026 · 10 citations
- Does Visual Pretraining Help End-to-End Reasoning?Chen Sun, Calvin Luo, Xingyi Zhou, Anurag Arnab et al.NeurIPS 2023 · 4 citations
- Look, Remember and Reason: Grounded Reasoning in Videos with Language ModelsApratim Bhattacharyya, Sunny Panchal, Reza Pourreza, Mingu Lee et al.ICLR 2024 · 15 citations
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen et al.ACM MM 2022 · 14 citations
- Flexible Context-Driven Sensory Processing in Dynamical Vision ModelsLakshmi Narasimhan Govindarajan, Abhiram Iyer, Valmiki Kothare, Ila FieteNeurIPS 2024 · 1 citation
