Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
Haotian Yan, Ming Wu, Chuang Zhang
Abstract
Multi-scale learning is central to semantic segmentation. We visualize the effective receptive field (ERF) of canonical multi-scale representations and point out two risks in learning them: scale inadequacy and field inactivation. A novel multi-scale learner, varying window attention (VWA), is presented to address these issues. VWA leverages the local window attention (LWA) and disentangles LWA into the query window and context window, allowing the context's scale to vary for the query to learn representations at multiple scales. However, varying the context to large-scale windows (enlarging ratio R) can significantly increase the memory footprint and computation cost (R^2 times larger than LWA). We propose a simple but professional re-scaling strategy to zero the extra induced cost without compromising performance. Consequently, VWA uses the same cost as LWA to overcome the receptive limitation of the local window. Furthermore, depending on VWA and employing various MLPs, we introduce a multi-scale decoder (MSD), VWFormer, to improve multi-scale representations for semantic segmentation. VWFormer achieves efficiency competitive with the most compute-friendly MSDs, like FPN and MLP decoder, but performs much better than any MSDs. For instance, using nearly half of UPerNet's computation, VWFormer outperforms it by 1.0%-2.5% mIoU on ADE20K. With little extra overhead, 10G FLOPs, Mask2Former armed with VWFormer improves by 1.0%-1.3%. The code and models are available at https://github.com/yan-hao-tian/vw
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentShi-Chen Zhang, Yunheng Li, Yu-Huan Wu, Qibin Hou et al.ICCV 2025 · 8 citations
- Differentiable Hierarchical Visual TokenizationMarius Aasan, Martine Hjelkrem-Tan, Nico Catalano, Changkyu Choi et al.NeurIPS 2025 · 4 citations
- NormDirection: Restoring the Missing Query Norm in Vision Linear AttentionWeikang Meng, Yadan Luo, Liangyu Huo, Yingjian Li et al.ICML 2026 · 2 citations
- Deeply Seeking Boundary for Lunar Regolith SegmentationYifeng Wang, Lingxin Wang, Lu Zhang, Yang Li et al.AAAI 2026
- WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image SegmentationSeunghun Moon, Hyunwoo Yu, Haeuk Lee, Suk-Ju KangICLR 2026
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
Related papers
- Dynamic Multi-Scale Filters for Semantic SegmentationJunjun He, Zhongying Deng, Yu QiaoICCV 2019 · 287 citations
- Dynamic Focus-aware Positional Queries for Semantic SegmentationHaoyu He, Jianfei Cai, Zizheng Pan, Jing Liu et al.CVPR 2023
- HRFormer: High-Resolution Vision Transformer for Dense PredictYuhui Yuan, Rao Fu, Lang Huang, Weihong Lin et al.NeurIPS 2021 · 357 citations
- MixFormer: Mixing Features across Windows and DimensionsQiang Chen, Qiman Wu, Jian Wang, Qinghao Hu et al.CVPR 2022 · 142 citations
- CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale AttentionWenxiao Wang, Lu Yao, Long Chen, Binbin Lin et al.ICLR 2022 · 367 citations
