Strip Pooling: Rethinking Spatial Pooling for Scene Parsing
Qibin Hou, Li Zhang, Ming-Ming Cheng, Jiashi Feng
Abstract
Spatial pooling has been proven highly effective in cap- turing long-range contextual information for pixel-wise prediction tasks, such as scene parsing. In this paper, beyond conventional spatial pooling that usually has a regular shape of N × N , we rethink the formulation of spatial pooling by introducing a new pooling strategy, called strip pooling, which considers a long but narrow kernel, i.e., 1 × N or N × 1. Based on strip pooling, we further investigate spatial pooling architecture design by 1) introducing a new strip pooling module that enables backbone networks to efficiently model long-range dependencies, 2) presenting a novel building block with diverse spatial pooling as a core, and 3) systematically comparing the performance of the proposed strip pooling and conventional spatial pooling techniques. Both novel pooling-based designs are lightweight and can serve as an efficient plugand-play module in existing scene parsing networks. Extensive experiments on popular benchmarks (e.g., ADE20K and Cityscapes) demonstrate that our simple approach establishes new state-of-the-art results. Code is available at https://github.com/Andrew-Qibin/SPNet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f8610fc-8d11-4675-ba03-a45106a1eca7Cited by top-tier papers37
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- All Tokens Matter: Token Labeling for Training Better Vision TransformersZihang Jiang, Qibin Hou, Li Yuan, Daquan Zhou et al.NeurIPS 2021 · 252 citations
- Representation Compensation Networks for Continual Semantic SegmentationChang-Bin Zhang, Jia-Wen Xiao, Xialei Liu, Ying-Cong Chen et al.CVPR 2022 · 102 citations
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationJiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß et al.CVPR 2022 · 100 citations
- AttaNet: Attention-Augmented Network for Fast and Accurate Scene ParsingQi Song, Kangfu Mei, Rui HuangAAAI 2021 · 89 citations
Builds on6
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- Boundary-Aware Feature Propagation for Scene SegmentationHenghui Ding, Xudong Jiang, Ai Qun Liu, Nadia Magnenat-Thalmann et al.ICCV 2019 · 283 citations
- SPGNet: Semantic Prediction Guidance for Scene ParsingBowen Cheng, Liang-Chieh Chen, Yunchao Wei, Yukun Zhu et al.ICCV 2019 · 117 citations
Related papers
- Adaptive Context Network for Scene ParsingJun Fu, Jing Liu, Yuhang Wang, Yong Li et al.ICCV 2019 · 148 citations
- Long Range Pooling for 3D Large-Scale Scene UnderstandingXiang-Li Li, Meng-Hao Guo, Tai-Jiang Mu, Ralph R. Martin et al.CVPR 2023
- Efficient Representation Learning via Adaptive Context PoolingChen Huang, Walter Talbott, Navdeep Jaitly, Joshua M. SusskindICML 2022 · 10 citations
- VSPW: A Large-scale Dataset for Video Scene Parsing in the WildJiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang et al.CVPR 2021
- Fully Attentional Network for Semantic SegmentationQi Song, Jie Li, Chenghong Li, Hao Guo et al.AAAI 2022 · 63 citations
