Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
Oleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki Pantazi
摘要
Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to outof-distribution visual inputs than prior models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- Emergent Symbols through Binding in External MemoryTaylor Whittington Webb, Ishan Sinha, Jonathan D. CohenICLR 2021 · 被引用 68 次
- Learning Representations that Support ExtrapolationTaylor W. Webb, Zachary Dulberg, Steven Frankland, Alexander A. Petrov 等ICML 2020 · 被引用 60 次
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
相关 Paper
- GAMR: A Guided Attention Model for (visual) ReasoningMohit Vaishnav, Thomas SerreICLR 2023 · 被引用 4 次
- ActiView: Evaluating Active Perception Ability for Multimodal Large Language ModelsZiyue Wang, Chi Chen, Fuwen Luo, Yurui Dong 等ACL 2025 · 被引用 9 次
- Progressive Visual Content Understanding Network for Image Emotion ClassificationJicai Pan, Shangfei WangACM MM 2023 · 被引用 5 次
- DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified ModelingXueyu Yuan, Jiarui Zhang, Jiangqi Song, Liu Liu 等ACM MM 2025
- REX: Reasoning-aware and Grounded ExplanationShi Chen, Qi ZhaoCVPR 2022 · 被引用 24 次
