Semantic Segmentation by Early Region Proxy
Yifan Zhang, Bo Pang, Cewu Lu
Abstract
Typical vision backbones manipulate structured features. As a compromise, semantic segmentation has long been modeled as per-point prediction on dense regular grids. In this work, we present a novel and efficient modeling that starts from interpreting the image as a tessellation of learnable regions, each of which has flexible geometrics and carries homogeneous semantics. To model region-wise context, we exploit Transformer to encode regions in a sequence-to-sequence manner by applying multi-layer self-attention on the region embeddings, which serve as proxies of specific regions. Semantic segmentation is now carried out as per-region prediction on top of the encoded region embeddings using a single linear classifier, where a decoder is no longer needed. The proposed RegProxy model discards the common Cartesian feature layout and operates purely at region level. Hence, it exhibits the most competitive performance-efficiency trade-off compared with the conventional dense prediction methods. For example, on ADE20K, the small-sized RegProxy-S/16 outperforms the best CNN model using 25% parameters and 4% computation, while the largest RegProxy-L/16 achieves 52.9 mIoU which outperforms the state-of-the-art by 2.1% with fewer resources. Codes and models are available at https://github.com/YiF-Zhang/RegionProxy .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Stochastic Segmentation with Conditional Categorical Diffusion ModelsLukas Zbinden, Lars Doorenbos, Theodoros Pissas, Adrian Thomas Huber et al.ICCV 2023 · 57 citations
- Focus on Query: Adversarial Mining Transformer for Few-Shot SegmentationYuan Wang, Naisong Luo, Tianzhu ZhangNeurIPS 2023 · 29 citations
- Learning Hierarchical Image Segmentation For Recognition and By RecognitionTsung-Wei Ke, Sangwoo Mo, Stella X. YuICLR 2024 · 20 citations
- Super-efficient Echocardiography Video Segmentation via Proxy- and Kernel-Based Semi-supervised LearningHuisi Wu, Jingyin Lin, Wende Xie, Jing QinAAAI 2023 · 16 citations
- Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic SegmentationYuan Wang, Rui Sun, Naisong Luo, Yuwen Pan et al.CVPR 2024 · 13 citations
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With TransformersSixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu et al.CVPR 2021
- PEM: Prototype-Based Efficient MaskFormer for Image SegmentationNiccolò Cavagnero, Gabriele Rosi, Claudia Cuttano, Francesca Pistilli et al.CVPR 2024
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Head-Free Lightweight Semantic Segmentation with Linear TransformerBo Dong, Pichao Wang, Fan WangAAAI 2023 · 129 citations
