No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition
Rong Qin, Xin Liu, Xingyu Liu, Jiaxuan Liu, Jinglei Shi, Liang Lin, Jufeng Yang
Abstract
Over the last decade, many notable methods have emerged to tackle the computational resource challenge of the high resolution image recognition (HRIR). They typically focus on identifying and aggregating a few salient regions for classification, discarding sub-salient areas for low training consumption. Nevertheless, many HRIR tasks necessitate the exploration of wider regions to model objects and contexts, which limits their performance in such scenarios. To address this issue, we present a DBPS strategy to enable training with more patches at low consumption. Specifically, in addition to a fundamental buffer that stores the embeddings of most salient patches, DBPS further employs an auxiliary buffer to recycle those sub-salient ones. To reduce the computational cost associated with gradients of sub-salient patches, these patches are primarily used in the forward pass to provide sufficient information for classification. Meanwhile, only the gradients of the salient patches are back-propagated to update the entire network. Moreover, we design a Multiple Instance Learning (MIL) architecture that leverages aggregated information from salient patches to filter out uninformative background within sub-salient patches for better accuracy. Besides, we introduce the random patch drop to accelerate training process and uncover informative regions. Experiment results demonstrate the superiority of our method in terms of both accuracy and training consumption against other advanced methods. The code is available in the https://github.com/Qinrong-NKU/DBPS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbb0a207-61fb-4397-8be6-07e685072760Cited by top-tier papers1
Ask how each one uses itBuilds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Multiple Instance Learning Framework with Masked Hard Instance Mining for Whole Slide Image ClassificationWenhao Tang, Sheng Huang, Xiaoxian Zhang, Fengtao Zhou et al.ICCV 2023 · 84 citations
- Feature Re-Embedding: Towards Foundation Model-Level Performance in Computational PathologyWenhao Tang, Fengtao Zhou, Sheng Huang, Xiang Zhu et al.CVPR 2024 · 70 citations
- Hard-Attention for Scalable Image ClassificationAthanasios Papadopoulos, Pawel Korus, Nasir D. MemonNeurIPS 2021 · 38 citations
Related papers
- Differentiable Patch Selection for Image RecognitionJean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn et al.CVPR 2021
- How Effective Can Dropout Be in Multiple Instance Learning ?Wenhui Zhu, Peijie Qiu, Xiwen Chen, Zhangsihao Yang et al.ICML 2025
- Iterative Patch Selection for High-Resolution Image RecognitionBenjamin Bergner, Christoph Lippert, Aravindh MahendranICLR 2023 · 4 citations
- PHR-DIFF: Portrait Highlights Removal via Patch-aware Diffusion ModelHongsheng Zheng, Zhongyun Bao, Gang Fu, Xuze Jiao et al.AAAI 2025 · 4 citations
- MIST: Multiple Instance Spatial TransformerBaptiste Angles, Yuhe Jin, Simon Kornblith, Andrea Tagliasacchi et al.CVPR 2021
