Consistency driven Sequential Transformers Attention Model for Partially Observable Scenes
Samrudhdhi B. Rangrej, Chetan L. Srinidhi, James J. Clark
Abstract
Most hard attention models initially observe a complete scene to locate and sense informative glimpses, and predict class-label of a scene based on glimpses. However, in many applications (e.g., aerial imaging), observing an entire scene is not always feasible due to the limited time and resources available for acquisition. In this paper, we develop a Sequential Transformers Attention Model (STAM) that only partially observes a complete image and predicts informative glimpse locations solely based on past glimpses. We design our agent using DeiT-distilled [44] and train it with a one-step actorcritic algorithm. Furthermore, to improve classification performance, we introduce a novel training objective, which enforces consistency between the class distribution predicted by a teacher model from a complete image and the class distribution predicted by our agent using glimpses. When the agent senses only 4% of the total image area, the inclusion of the proposed consistency loss in our training objective yields 3% and 8% higher accuracy on ImageNet and fMoW datasets, respectively. Moreover, our agent outperforms previous state-of-the-art by observing nearly 27% and 42% fewer pixels in glimpses on ImageNet and fMoW.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cfb3210-02f7-4dc5-8b46-128b066146a3Cited by top-tier papers5
- Internally Rewarded Reinforcement LearningMengdi Li, Xufeng Zhao, Jae Hee Lee, Cornelius Weber et al.ICML 2023 · 17 citations
- Online Feedback Efficient Active Target Discovery in Partially Observable EnvironmentsAnindya Sarkar, Binglin Ji, Yevgeniy VorobeychikNeurIPS 2025 · 1 citation
- Masked Jigsaw Puzzle: A Versatile Position Embedding for Vision TransformersBin Ren, Yahui Liu, Yue Song, Wei Bi et al.CVPR 2023
- Active Target Discovery under Uninformative Priors: The Power of Permanent and Transient MemoryAnindya Sarkar, Binglin Ji, Yevgeniy VorobeychikNeurIPS 2025
- InCoDe: Interpretable Compressed Descriptions For Image GenerationArmand Comas Massague, Aditya Chattopadhyay, Feliu Formosa, Changyu Liu et al.ICLR 2025
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
Related papers
- Transformer-based Open-world Instance Segmentation with Cross-task Consistency RegularizationXizhe Xue, Dongdong Yu, Lingqiao Liu, Yu Liu et al.ACM MM 2023 · 2 citations
- RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionYunqing Hu, Xuan Jin, Yin Zhang, Haiwen Hong et al.ACM MM 2021 · 142 citations
- LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Changan Wang, Yabiao Wang, Guannan Jiang et al.AAAI 2022 · 61 citations
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou et al.NeurIPS 2021 · 252 citations
- Partial Class Activation Attention for Semantic SegmentationSun'ao Liu, Hongtao Xie, Hai Xu, Yongdong Zhang et al.CVPR 2022 · 47 citations
