Hard-Attention for Scalable Image Classification
Athanasios Papadopoulos, Pawel Korus, Nasir D. Memon
Abstract
Can we leverage high-resolution information without the unsustainable quadratic complexity to input scale? We propose Traversal Network (TNet), a novel multi-scale hard-attention architecture, which traverses image scale-space in a top-down fashion, visiting only the most informative image regions along the way. TNet offers an adjustable trade-off between accuracy and complexity, by changing the number of attended image locations. We compare our model against hard-attention baselines on ImageNet, achieving higher accuracy with less resources (FLOPs, processing time and memory). We further test our model on fMoW dataset, where we process satellite images of size up to px, getting up to x faster processing compared to baselines operating on the same resolution, while achieving higher accuracy as well. TNet is modular, meaning that most classification models could be adopted as its backbone for feature extraction, making the reported performance gains orthogonal to benefits offered by existing optimized deep models. Finally, hard-attention guarantees a degree of interpretability to our model's predictions, without any extra cost beyond inference. Code is available at .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f41736e0-c072-4a54-b459-af30ffb6a6d3Cited by top-tier papers5
- Efficient Classification of Very Large Images with Tiny ObjectsFanjie Kong, Ricardo HenaoCVPR 2022 · 32 citations
- A Partially-Supervised Reinforcement Learning Framework for Visual Active SearchAnindya Sarkar, Nathan Jacobs, Yevgeniy VorobeychikNeurIPS 2023 · 13 citations
- Consistency driven Sequential Transformers Attention Model for Partially Observable ScenesSamrudhdhi B. Rangrej, Chetan L. Srinidhi, James J. ClarkCVPR 2022 · 8 citations
- No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image RecognitionRong Qin, Xin Liu, Xingyu Liu, Jiaxuan Liu et al.CVPR 2025
- Learning Dynamics Feature Representation via Policy Attention for Dynamic Path Planning in Urban Road NetworksKai Zhang, Jingjing Gu, Qiuhong WangICLR 2026
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- Learning Attentive Pairwise Interaction for Fine-Grained ClassificationPeiqin Zhuang, Yali Wang, Yu QiaoAAAI 2020 · 392 citations
Related papers
- Global Filter Networks for Image ClassificationYongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu et al.NeurIPS 2021 · 798 citations
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- Quadtree Attention for Vision TransformersShitao Tang, Jiahui Zhang, Siyu Zhu, Ping TanICLR 2022 · 194 citations
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock et al.ICCV 2021 · 1,009 citations
