Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoning
Oleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki Pantazi
Abstract
Human capabilities in understanding visual relations are far superior to those of AI systems, especially for previously unseen objects. For example, while AI systems struggle to determine whether two such objects are visually the same or different, humans can do so with ease. Active vision theories postulate that the learning of visual relations is grounded in actions that we take to fixate objects and their parts by moving our eyes. In particular, the low-dimensional spatial information about the corresponding eye movements is hypothesized to facilitate the representation of relations between different image parts. Inspired by these theories, we develop a system equipped with a novel Glimpse-based Active Perception (GAP) that sequentially glimpses at the most salient regions of the input image and processes them at high resolution. Importantly, our system leverages the locations stemming from the glimpsing actions, along with the visual content around them, to represent relations between different parts of the image. The results suggest that the GAP is essential for extracting visual relations that go beyond the immediate visual content. Our approach reaches state-of-the-art performance on several visual reasoning tasks being more sample-efficient, and generalizing better to outof-distribution visual inputs than prior models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80525975-8e29-458b-bd56-8d15221aa09aBuilds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Emergent Symbols through Binding in External MemoryTaylor Whittington Webb, Ishan Sinha, Jonathan D. CohenICLR 2021 · 68 citations
- Learning Representations that Support ExtrapolationTaylor W. Webb, Zachary Dulberg, Steven Frankland, Alexander A. Petrov et al.ICML 2020 · 60 citations
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 35 citations
Related papers
- GAMR: A Guided Attention Model for (visual) ReasoningMohit Vaishnav, Thomas SerreICLR 2023 · 4 citations
- ActiView: Evaluating Active Perception Ability for Multimodal Large Language ModelsZiyue Wang, Chi Chen, Fuwen Luo, Yurui Dong et al.ACL 2025 · 9 citations
- Progressive Visual Content Understanding Network for Image Emotion ClassificationJicai Pan, Shangfei WangACM MM 2023 · 5 citations
- DFGAP: Towards Depth-Free Cross-Category GAParts Perception via Uncertainty-Quantified ModelingXueyu Yuan, Jiarui Zhang, Jiangqi Song, Liu Liu et al.ACM MM 2025
- REX: Reasoning-aware and Grounded ExplanationShi Chen, Qi ZhaoCVPR 2022 · 24 citations
