FeedFormer: Revisiting Transformer Decoder for Efficient Semantic Segmentation
Jae-hun Shim, Hyunwoo Yu, Kyeongbo Kong, Suk-Ju Kang
摘要
With the success of Vision Transformer (ViT) in image classification, its variants have yielded great success in many downstream vision tasks. Among those, the semantic segmentation task has also benefited greatly from the advance of ViT variants. However, most studies of the transformer for semantic segmentation only focus on designing efficient transformer encoders, rarely giving attention to designing the decoder. Several studies make attempts in using the transformer decoder as the segmentation decoder with class-wise learnable query. Instead, we aim to directly use the encoder features as the queries. This paper proposes the Feature Enhancing Decoder transFormer (FeedFormer) that enhances structural information using the transformer decoder. Our goal is to decode the high-level encoder features using the lowest-level encoder feature. We do this by formulating high-level features as queries, and the lowest-level feature as the key and value. This enhances the high-level features by collecting the structural information from the lowest-level feature. Additionally, we use a simple reformation trick of pushing the encoder blocks to take the place of the existing self-attention module of the decoder to improve efficiency. We show the superiority of our decoder with various light-weight transformer-based decoders on popular semantic segmentation datasets. Despite the minute computation, our model has achieved state-of-the-art performance in the performance computation trade-off. Our model FeedFormer-B0 surpasses SegFormer-B0 with 1.8% higher mIoU and 7.1% less computation on ADE20K, and 1.7% higher mIoU and 14.4% less computation on Cityscapes, respectively. Code will be released at: https://github.com/jhshim1995/FeedFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Generalized Medical Image Segmentation from Decoupled Feature QueriesQi Bi, Jingjun Yi, Hao Zheng, Wei Ji 等AAAI 2024 · 被引用 41 次
- MaSS13K: A Matting-level Semantic Segmentation BenchmarkChenxi Xie, Minghan Li, Hui Zeng, Jun Luo 等CVPR 2025
- WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image SegmentationSeunghun Moon, Hyunwoo Yu, Haeuk Lee, Suk-Ju KangICLR 2026
- Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building LabelsJunda Xu, Yanmeng Liu, Xiangqiang Zeng, Jinrong Wu 等CVPR 2026
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang 等ICCV 2019 · 被引用 2,972 次
相关 Paper
- Head-Free Lightweight Semantic Segmentation with Linear TransformerBo Dong, Pichao Wang, Fan WangAAAI 2023 · 被引用 129 次
- SegViT: Semantic Segmentation with Plain Vision TransformersBowen Zhang, Zhi Tian, Quan Tang, Xiangxiang Chu 等NeurIPS 2022 · 被引用 242 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
- MP-Former: Mask-Piloted Transformer for Image SegmentationHao Zhang, Feng Li, Huaizhe Xu, Shijia Huang 等CVPR 2023
- InterFormer Real-time Interactive Image SegmentationYou Huang, Hao Yang, Ke Sun, Shengchuan Zhang 等ICCV 2023 · 被引用 36 次
