WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation
Seunghun Moon, Hyunwoo Yu, Haeuk Lee, Suk-Ju Kang
Abstract
Existing Parameter-Efficient Tuning (PET) methods for Referring Image Segmentation (RIS) primarily focus on layer-wise feature alignment, often neglecting the crucial role of a neck module for the intermediate fusion of aggregated multi-scale features, which creates a significant performance bottleneck. To address this limitation, we introduce WIMFRIS, a novel framework that establishes a powerful neck architecture alongside a simple yet effective PET strategy. At its core is our proposed Hierarchical Mamba Fusion (HMF) block, which first aggregates multi-scale features and then employs a novel Window Mamba Fuser (WMF) module to perform effective intermediate fusion. This WMF module leverages non-overlapping window partitioning to mitigate the information decay problem inherent in State-Space Models (SSMs) while ensuring rich local-global context interaction. Furthermore, our PET strategy enhances primary alignment with a Mamba Text Adapter (MTA) for robust textual priors, a Multi-Scale Aligner (MSA) for precise vision-language fusion, and learnable emphasis parameters for adaptive stage-wise feature weighting. Extensive experiments demonstrate that WIMFRIS achieves new state-of-the-art performance across all public RIS benchmarks. Code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3cf3e4e-27c2-444f-a10e-3f27ea524dd3Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi et al.ICLR 2024 · 11 citations
- Adaptive Selection based Referring Image SegmentationPengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu et al.ACM MM 2024 · 18 citations
- Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationZunnan Xu, Zhihong Chen, Yong Zhang, Yibing Song et al.ICCV 2023 · 85 citations
- Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationJiaqi Huang, Zunnan Xu, Ting Liu, Yong Liu et al.AAAI 2025 · 34 citations
- Deep Instruction Tuning for Segment Anything ModelXiaorui Huang, Gen Luo, Chaoyang Zhu, Bo Tong et al.ACM MM 2024 · 3 citations
