LIP: Local Importance-Based Pooling
Ziteng Gao, Limin Wang, Gangshan Wu
Abstract
Spatial downsampling layers are favored in convolutional neural networks (CNNs) to downscale feature maps for larger receptive fields and less memory consumption. However, for visual recognition tasks, these layers might lose discriminative details due to improper pooling strategies. In this paper, we present a unified framework (LAN) over the common downsampling layers (e.g., average pooling, max pooling, and strided convolution) from a view of local aggregation based on importance. In this LAN framework, we analyze the issues of these widely-used pooling layers and figure out the criteria of designing an effective downsampling layer. Based on this analysis, we propose a simple, general, and effective pooling operation based on local importance modeling, termed as Local Importance-based Pooling (LIP). LIP is able to enhance discriminative features during the downsampling procedure by learning adaptive importance weights based on inputs. To further modulate different pooling windows for more effective pooling, we present the improved version of LIP, termed LIP++, by introducing an explicit margin term and efficient logit modules. Our LIP++ can yield consistent accuracy improvement over the original LIP yet with a smaller computational cost. Extensive experiments show that our presented LIP method consistently yields notable gains with different CNN architectures on the image classification task. In the challenging MS COCO dataset, detectors with our LIP-ResNets as backbones obtain a consistent performance improvement over the vanilla ResNets on both bounding box detection and instance segmentation. Finally, we also verify the effectiveness of LIP on the tasks of pose estimation and semantic segmentation, demonstrating its generalization to the dense prediction task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39302e3c-0e01-43e8-bfa1-1fb914178db5Cited by top-tier papers6
- Refining activation downsampling with SoftPoolAlexandros Stergiou, Ronald Poppe, Grigorios KalliatakisICCV 2021 · 195 citations
- HiFaceGAN: Face Renovation via Collaborative Suppression and ReplenishmentLingbo Yang, Shanshe Wang, Siwei Ma, Wen Gao et al.ACM MM 2020 · 136 citations
- Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object DetectionXiaonan Lu, Wenhui Diao, Yongqiang Mao, Junxi Li et al.AAAI 2023 · 66 citations
- DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid IntegrationWeiran Chen, Guiqian Zhu, Ying Li, Yi Ji et al.ACM MM 2025 · 2 citations
- CAPE: CAM as a Probabilistic Ensemble for Enhanced DNN InterpretationTownim Faisal Chowdhury, Kewen Liao, Vu Minh Hieu Phan, Minh-Son To et al.CVPR 2024
Builds on2
Related papers
- Global Feature Guided Local PoolingTakumi KobayashiICCV 2019 · 24 citations
- Information Entropy Based Feature Pooling for Convolutional Neural NetworksWeitao Wan, Jiansheng Chen, Tianpeng Li, Yiqing Huang et al.ICCV 2019 · 33 citations
- D2Det: Towards High Quality Object Detection and Instance SegmentationJiale Cao, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan et al.CVPR 2020
- What Deep CNNs Benefit From Global Covariance Pooling: An Optimization PerspectiveQilong Wang, Li Zhang, Banggu Wu, Dongwei Ren et al.CVPR 2020
- Context-aware Attentional Pooling (CAP) for Fine-grained Visual ClassificationArdhendu Behera, Zachary Wharton, Pradeep R. P. G. Hewage, Asish BeraAAAI 2021 · 142 citations
