Efficient Classification of Very Large Images with Tiny Objects
Fanjie Kong, Ricardo Henao
Abstract
An increasing number of applications in computer vision, specially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b60e9d8-9cef-4767-b5bd-2c69eae70a5fCited by top-tier papers3
- Few-Shot Class-Incremental Learning for Named Entity RecognitionRui Wang, Tong Yu, Handong Zhao, Sungchul Kim et al.ACL 2022 · 26 citations
- Visual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology ImagesMing Y. Lu, Bowen Chen, Andrew Zhang, Drew F. K. Williamson et al.CVPR 2023
- No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image RecognitionRong Qin, Xin Liu, Xingyu Liu, Jiaxuan Liu et al.CVPR 2025
Builds on4
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Hard-Attention for Scalable Image ClassificationAthanasios Papadopoulos, Pawel Korus, Nasir D. MemonNeurIPS 2021 · 38 citations
- Learning When and Where to Zoom With Deep Reinforcement LearningBurak Uzkent, Stefano ErmonCVPR 2020
- Differentiable Patch Selection for Image RecognitionJean-Baptiste Cordonnier, Aravindh Mahendran, Alexey Dosovitskiy, Dirk Weissenborn et al.CVPR 2021
Related papers
- Sequential Attention-based Sampling for Histopathological AnalysisTarun Gogisetty, Naman Malpani, Gugan Thoppe, Sridharan DevarajanNeurIPS 2025 · 2 citations
- Multi-Stage Pathological Image Classification Using Semantic SegmentationShusuke Takahama, Yusuke Kurose, Yusuke Mukuta, Hiroyuki Abe et al.ICCV 2019 · 53 citations
- ThumbNet: One Thumbnail Image Contains All You Need for RecognitionChen Zhao, Bernard GhanemACM MM 2020 · 14 citations
- Recurrent Networks for Guided Multi-Attention ClassificationXin Dai, Xiangnan Kong, Tian Guo, John Boaz Lee et al.KDD 2020 · 5 citations
- Bridging Local Inductive Bias and Long-Range Dependencies With Pixel-Mamba for End-To-End Whole Slide Image AnalysisZhongwei Qiu, Hanqing Chao, Tiancheng Lin, Wanxing Chang et al.ICCV 2025 · 1 citation
