SegMAN: Omni-scale Context Modeling with State Space Models and Local Attention for Semantic Segmentation
Yunxiang Fu, Meng Lou, Yizhou Yu
Abstract
High-quality semantic segmentation relies on three key capabilities: global context modeling, local detail encoding, and multi-scale feature extraction. However, recent methods struggle to possess all these capabilities simultaneously. Hence, we aim to empower segmentation networks to simultaneously carry out efficient global context modeling, high-quality local detail encoding, and rich multi-scale feature representation for varying input resolutions. In this paper, we introduce SegMAN, a novel linear-time model comprising a hybrid feature encoder dubbed SegMAN Encoder, and a decoder based on state space models. Specifically, the SegMAN Encoder synergistically integrates sliding local attention with dynamic state space models, enabling highly efficient global context modeling while preserving fine-grained local details. Meanwhile, the MMSCopE module in our decoder enhances multi-scale context feature extraction and adaptively scales with the input resolution. Our SegMAN-B Encoder achieves 85.1% ImageNet-1k accuracy (+1.5% over VMamba-S with fewer parameters). When paired with our decoder, the full SegMAN-B model achieves 52.6% mIoU on ADE20K (+1.6% over SegNeXt-L with 15% fewer GFLOPs), 83.8% mIoU on Cityscapes (+2.1% over SegFormer-B3 with half the GFLOPs), and 1.6% higher mIoU than VWFormer-B3 on COCO-Stuff with lower GFLOPs. Our code is available at https:// github.com/yunxiangfu2001/SegMAN .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e524e8ae-2f54-41a6-bd33-0c5ffd30f4a1Cited by top-tier papers10
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 17 citations
- HydraMamba: Multi-Head State Space Model for Global Point Cloud LearningKanglin Qu, Pan Gao, Qun Dai, Yuanhao SunACM MM 2025 · 2 citations
- SLARM: Streaming and Language-Aligned Reconstruction Model for Dynamic ScenesZhicheng Qiu, Jiarui Meng, Tong-an Luo, Yican Huang et al.CVPR 2026 · 2 citations
- XSeg: A Large-scale X-ray Contraband Segmentation Benchmark For Real-World Security ScreeningHongxia Gao, Yixin Chen, Jiali Wen, Litao Li et al.CVPR 2026 · 1 citation
- Deeply Seeking Boundary for Lunar Regolith SegmentationYifeng Wang, Lingxin Wang, Lu Zhang, Yang Li et al.AAAI 2026
Builds on32
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- Efficient Parallel Multi-Scale Detail and Semantic Encoding Network for Lightweight Semantic SegmentationXiao Liu, Xiuya Shi, Lufei Chen, Linbo Qing et al.ACM MM 2023 · 7 citations
- AttaNet: Attention-Augmented Network for Fast and Accurate Scene ParsingQi Song, Kangfu Mei, Rui HuangAAAI 2021 · 89 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- Dynamic Multi-Scale Filters for Semantic SegmentationJunjun He, Zhongying Deng, Yu QiaoICCV 2019 · 287 citations
- RTFormer: Efficient Design for Real-Time Semantic Segmentation with TransformerJian Wang, Chenhui Gou, Qiman Wu, Haocheng Feng et al.NeurIPS 2022 · 207 citations
