GAMR: A Guided Attention Model for (visual) Reasoning
Mohit Vaishnav, Thomas Serre
摘要
Humans continue to outperform modern AI systems in their ability to flexibly parse and understand complex visual scenes. Here, we present a novel module for visual reasoning, the Guided Attention Model for (visual) Reasoning (GAMR), which instantiates an active vision theory -- positing that the brain solves complex visual reasoning problems dynamically -- via sequences of attention shifts to select and route task-relevant visual information into memory. Experiments on an array of visual reasoning tasks and datasets demonstrate GAMR's ability to learn visual routines in a robust and sample-efficient manner. In addition, GAMR is shown to be capable of zero-shot generalization on completely novel reasoning tasks. Overall, our work provides computational support for cognitive theories that postulate the need for a critical interplay between attention and memory to dynamically maintain and manipulate task-relevant visual information to solve complex visual reasoning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Large Language Models are Visual Reasoning CoordinatorsLiangyu Chen, Bo Li, Sheng Shen, Jingkang Yang 等NeurIPS 2023 · 被引用 108 次
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata 等NeurIPS 2024 · 被引用 101 次
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 被引用 10 次
- Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosShantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston TanNeurIPS 2024 · 被引用 9 次
它引用的顶会 Paper5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
- Stratified Rule-Aware Network for Abstract Visual ReasoningSheng Hu, Yuqing Ma, Xianglong Liu, Yanlu Wei 等AAAI 2021 · 被引用 126 次
- Emergent Symbols through Binding in External MemoryTaylor Whittington Webb, Ishan Sinha, Jonathan D. CohenICLR 2021 · 被引用 68 次
相关 Paper
- Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual ReasoningFan Shi, Bin Li, Xiangyang XueICML 2025
- Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal ReasoningYang Zhang, Xiaoshuai Sun, Rui Zhao, Wujin Sun 等ICML 2026
- Attention over Learned Object Embeddings Enables Complex Visual ReasoningDavid Ding, Felix Hill, Adam Santoro, Malcolm Reynolds 等NeurIPS 2021 · 被引用 87 次
- Grounded Reinforcement Learning for Visual ReasoningGabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain 等NeurIPS 2025 · 被引用 90 次
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual ReasoningZejun Li, Yingxiu Zhao, Jiwen Zhang, Siyuan Wang 等ICLR 2026 · 被引用 7 次
