Differentiable Patch Selection for Image Recognition
Jean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn, Jakob Uszkoreit, Thomas Unterthiner
Abstract
Neural Networks require large amounts of memory and compute to process high resolution images, even when only a small part of the image is actually informative for the task at hand. We propose a method based on a differentiable Top-K operator to select the most relevant parts of the input to efficiently process high resolution images. Our method may be interfaced with any downstream neural network, is able to aggregate information from different patches in a flexible way, and allows the whole model to be trained endto-end using backpropagation. We show results for traffic sign recognition, inter-patch relationship reasoning, and fine-grained recognition without using object/part bounding box annotations during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a83c8c50-4cd2-4a94-9baa-14972f92ace8Cited by top-tier papers30
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 281 citations
- Towards mental time travel: a hierarchical memory for reinforcement learning agentsAndrew K. Lampinen, Stephanie C. Y. Chan, Andrea Banino, Felix HillNeurIPS 2021 · 63 citations
- Differentiable Top-k Classification LearningFelix Petersen, Hilde Kuehne, Christian Borgelt, Oliver DeussenICML 2022 · 48 citations
- Fast, Differentiable and Sparse Top-k: a Convex Analysis PerspectiveMichael Eli Sander, Joan Puigcerver, Josip Djolonga, Gabriel Peyré et al.ICML 2023 · 35 citations
- KVQ: Kwai Video Quality Assessment for Short-form VideosYiting Lu, Xin Li, Yajing Pei, Kun Yuan et al.CVPR 2024 · 32 citations
Builds on3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 285 citations
Related papers
- Iterative Patch Selection for High-Resolution Image RecognitionBenjamin Bergner, Christoph Lippert, Aravindh MahendranICLR 2023 · 4 citations
- MIST: Multiple Instance Spatial TransformerBaptiste Angles, Yuhe Jin, Simon Kornblith, Andrea Tagliasacchi et al.CVPR 2021
- Multi-Stage Pathological Image Classification Using Semantic SegmentationShusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe et al.ICCV 2019 · 53 citations
- Differentiable Hierarchical Visual TokenizationMarius Aasan, Martine Hjelkrem-Tan, Nico Catalano, Changkyu Choi et al.NeurIPS 2025 · 4 citations
- Learning Compositional Neural Information Fusion for Human ParsingWenguan Wang, Zhijie Zhang, Siyuan Qi, Jianbing Shen et al.ICCV 2019 · 131 citations
