Iterative Patch Selection for High-Resolution Image Recognition
Benjamin Bergner, Christoph Lippert, Aravindh Mahendran
Abstract
High-resolution images are prevalent in various applications, such as autonomous driving and computer-aided diagnosis. However, training neural networks on such images is computationally challenging and easily leads to out-of-memory errors even on modern GPUs. We propose a simple method, Iterative Patch Selection (IPS), which decouples the memory usage from the input size and thus enables the processing of arbitrarily large images under tight hardware constraints. IPS achieves this by selecting only the most salient patches, which are then aggregated into a global representation for image recognition. For both patch selection and aggregation, a cross-attention based transformer is introduced, which exhibits a close connection to Multiple Instance Learning. Our method demonstrates strong performance and has wide applicability across different domains, training regimes and image sizes while using minimal accelerator memory. For example, we are able to finetune our model on whole-slide images consisting of up to 250k patches (>16 gigapixels) with only 5 GB of GPU VRAM at a batch size of 16.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisHonglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui et al.NeurIPS 2024 · 27 citations
- Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyWenhao Tang, Rong Qin, Heng Fang, Fengtao Zhou et al.NeurIPS 2025 · 10 citations
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-SupervisionAnthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim et al.NeurIPS 2025 · 7 citations
- PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationHonglin Li, Zhongyi Shui, Yunlong Zhang, Chenglu Zhu et al.NeurIPS 2025 · 6 citations
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionQiyu Xu, Zhanxuan Hu, Yu Duan, Ercheng Pei et al.ICCV 2025 · 5 citations
Builds on9
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
Related papers
- Differentiable Patch Selection for Image RecognitionJean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn et al.CVPR 2021
- Adaptive Patching for High-resolution Image Segmentation with TransformersEnzhi Zhang, Isaac Lyngaas, Peng Chen, Xiao Wang et al.SC 2024 · 6 citations
- Multi-Stage Pathological Image Classification Using Semantic SegmentationShusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe et al.ICCV 2019 · 53 citations
- Rotation-Agnostic Image Representation Learning for Digital PathologySaghir Alfasly, Abubakr Shafique, Peyman Nejat, Jibran A. Khan et al.CVPR 2024
- Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationYingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning WangAAAI 2024 · 10 citations
