Prompting Multi-Modal Image Segmentation with Semantic Grouping
Qibin He
Abstract
Multi-modal image segmentation is one of the core issues in computer vision. The main challenge lies in integrating common information between modalities while retaining specific patterns for each modality. Existing methods typically perform full fine-tuning on RGB-based pre-trained parameters to inherit the powerful representation of the foundation model. Although effective, such paradigm is not optimal due to weak transferability and scarce downstream data. Inspired by the recent success of prompt learning in language models, we propose the Grouping Prompt Tuning Framework (GoPT), which introduces explicit semantic grouping to learn modal-related prompts, adapting the frozen pre-trained foundation model to various downstream multi-modal segmentation tasks. Specifically, a class-aware uni-modal prompter is designed to balance intra- and inter-modal semantic propagation by grouping modality-specific class tokens, thereby improving the adaptability of spatial information. Furthermore, an alignment-induced cross-modal prompter is introduced to aggregate class-aware representations and share prompt parameters among different modalities to assist in modeling common statistics. Extensive experiments show the superiority of our GoPT, which achieves SOTA performance on various downstream multi-modal image segmentation tasks by training only < 1% model parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc265382-4e0e-457c-a57f-97deb16217afCited by top-tier papers5
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationBingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao et al.ACM MM 2025 · 15 citations
- OGP-Net: Optical Guidance Meets Pixel-Level Contrastive Distillation for Robust Multi-Modal and Missing Modality SegmentationAniruddh Sikdar, Jayant Teotia, Suresh SundaramAAAI 2025 · 8 citations
- MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsDongpan Chen, Dehui Kong, Jinghua Li, Baocai YinAAAI 2025 · 5 citations
- Decoupled and Reusable Adaptation for Efficient Cross-Modal TransferYajing Liu, Yumeng Zhang, Yue Si, Baojie Fan et al.CVPR 2026
- Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic SegmentationJiaxin Cai, Jingze Su, Qi Li, Wenjie Yang et al.CVPR 2025
Builds on9
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- Edge-Aware Guidance Fusion Network for RGB-Thermal Scene ParsingWujie Zhou, Shaohua Dong, Caie Xu, Yaguan QianAAAI 2022 · 151 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer FusionYikai Wang, Fuchun Sun, Ming Lu, Anbang YaoACM MM 2020 · 66 citations
- Fine-tuning Image Transformers using Learnable MemoryMark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew JacksonCVPR 2022 · 51 citations
Related papers
- Visual Prompt Multi-Modal TrackingJiawen Zhu, Simiao Lai, Xin Chen, Dong Wang et al.CVPR 2023
- MixPrompt: Efficient Mixed Prompting for Multimodal Semantic SegmentationZhiwei Hao, Zhongyu Xiao, Jianyuan Guo, Li Shen et al.NeurIPS 2025 · 1 citation
- Tuning Multi-mode Token-level Prompt Alignment across ModalitiesDongsheng Wang, Miaoge Li, Xinyang Liu, Mingsheng Xu et al.NeurIPS 2023 · 49 citations
- X-Prompt: Multi-modal Visual Prompt for Video Object SegmentationPinxue Guo, Wanyun Li, Hao Huang, Lingyi Hong et al.ACM MM 2024 · 7 citations
- Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt TuningLei-Lei Ma, Shuo Xu, Ming-Kun Xie, Lei Wang et al.CVPR 2025
