Iterative Patch Selection for High-Resolution Image Recognition
Benjamin Bergner, Christoph Lippert, Aravindh Mahendran
摘要
High-resolution images are prevalent in various applications, such as autonomous driving and computer-aided diagnosis. However, training neural networks on such images is computationally challenging and easily leads to out-of-memory errors even on modern GPUs. We propose a simple method, Iterative Patch Selection (IPS), which decouples the memory usage from the input size and thus enables the processing of arbitrarily large images under tight hardware constraints. IPS achieves this by selecting only the most salient patches, which are then aggregated into a global representation for image recognition. For both patch selection and aggregation, a cross-attention based transformer is introduced, which exhibits a close connection to Multiple Instance Learning. Our method demonstrates strong performance and has wide applicability across different domains, training regimes and image sizes while using minimal accelerator memory. For example, we are able to finetune our model on whole-slide images consisting of up to 250k patches (>16 gigapixels) with only 5 GB of GPU VRAM at a batch size of 16.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisHonglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui 等NeurIPS 2024 · 被引用 27 次
- Revisiting End-to-End Learning with Slide-level Supervision in Computational PathologyWenhao Tang, Rong Qin, Heng Fang, Fengtao Zhou 等NeurIPS 2025 · 被引用 10 次
- LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-SupervisionAnthony Fuller, Yousef Yassin, Junfeng Wen, Tarek Ibrahim 等NeurIPS 2025 · 被引用 7 次
- PathVQ: Reforming Computational Pathology Foundation Model for Whole Slide Image Analysis via Vector QuantizationHonglin Li, Zhongyi Shui, Yunlong Zhang, Chenglu Zhu 等NeurIPS 2025 · 被引用 6 次
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionQiyu Xu, Zhanxuan Hu, Yu Duan, Ercheng Pei 等ICCV 2025 · 被引用 5 次
它引用的顶会 Paper9
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen 等CVPR 2022 · 被引用 490 次
相关 Paper
- Differentiable Patch Selection for Image RecognitionJean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn 等CVPR 2021
- Adaptive Patching for High-resolution Image Segmentation with TransformersEnzhi Zhang, Isaac Lyngaas, Peng Chen, Xiao Wang 等SC 2024 · 被引用 6 次
- Multi-Stage Pathological Image Classification Using Semantic SegmentationShusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe 等ICCV 2019 · 被引用 53 次
- Rotation-Agnostic Image Representation Learning for Digital PathologySaghir Alfasly, Abubakr Shafique, Peyman Nejat, Jibran A. Khan 等CVPR 2024
- Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationYingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning WangAAAI 2024 · 被引用 10 次
