WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation
Seunghun Moon, Hyunwoo Yu, Haeuk Lee, Suk-Ju Kang
摘要
Existing Parameter-Efficient Tuning (PET) methods for Referring Image Segmentation (RIS) primarily focus on layer-wise feature alignment, often neglecting the crucial role of a neck module for the intermediate fusion of aggregated multi-scale features, which creates a significant performance bottleneck. To address this limitation, we introduce WIMFRIS, a novel framework that establishes a powerful neck architecture alongside a simple yet effective PET strategy. At its core is our proposed Hierarchical Mamba Fusion (HMF) block, which first aggregates multi-scale features and then employs a novel Window Mamba Fuser (WMF) module to perform effective intermediate fusion. This WMF module leverages non-overlapping window partitioning to mitigate the information decay problem inherent in State-Space Models (SSMs) while ensuring rich local-global context interaction. Furthermore, our PET strategy enhances primary alignment with a Mamba Text Adapter (MTA) for robust textual priors, a Multi-Scale Aligner (MSA) for precise vision-language fusion, and learnable emphasis parameters for adaptive stage-wise feature weighting. Extensive experiments demonstrate that WIMFRIS achieves new state-of-the-art performance across all public RIS benchmarks. Code is available here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- BarLeRIa: An Efficient Tuning Framework for Referring Image SegmentationYaoming Wang, Jin Li, Xiaopeng Zhang, Bowen Shi 等ICLR 2024 · 被引用 11 次
- Adaptive Selection based Referring Image SegmentationPengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu 等ACM MM 2024 · 被引用 18 次
- Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image SegmentationZunnan Xu, Zhihong Chen, Yong Zhang, Yibing Song 等ICCV 2023 · 被引用 85 次
- Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationJiaqi Huang, Zunnan Xu, Ting Liu, Yong Liu 等AAAI 2025 · 被引用 34 次
- Deep Instruction Tuning for Segment Anything ModelXiaorui Huang, Gen Luo, Chaoyang Zhu, Bo Tong 等ACM MM 2024 · 被引用 3 次
