MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot Segmentation
Hang Xiong, Runmin Cong, Jinpeng Chen, Chen Zhang, Feng Li, Huihui Bai, Sam Kwong
Abstract
Despite the effectiveness of Segment Anything Model (SAM) based methods in Few-Shot Segmentation (FSS) tasks, our closer examination of their prompt encoding mechanism reveals that these methods rely solely on visual information to generate a single type of prompt. Consequently, they suffer from semantic granularity representation bias and a loss of spatial information. To address these limitations, this paper introduces an innovative multi-modal prompt encoder, enabling SAM to leverage both annotated reference images and textual descriptions of class names as segmentation prompts. This approach generates text prompts, dense visual prompts, and sparse visual prompts, spanning multiple modalities and granularities. These prompts provide enhanced representations of the target class, capturing both abstract semantics and specific details, while ensuring granularity appropriateness. When our multi-modal prompt encoder is integrated with SAM's image encoder and mask decoder, the overall model is referred to as MM-Prompt. To validate its effectiveness, we conducted extensive empirical studies on the PASCAL-5^i and COCO-20^i datasets. The experimental results demonstrate that MM-Prompt achieves state-of-the-art performance in FSS tasks, highlighting its substantial potential and value in this domain.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0bff1754-c8be-4e5c-9a1e-f88cc0a20bbbCited by top-tier papers1
Ask how each one uses itRelated papers
- VRP-SAM: SAM with Visual Reference PromptYanpeng Sun, Jiahui Chen, Shan Zhang, Xinyu Zhang et al.CVPR 2024 · 49 citations
- APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic SegmentationWeizhao He, Yang Zhang, Wei Zhuo, Linlin Shen et al.CVPR 2024
- Focus on Background: Exploring SAM's Potential in Few-shot Medical Image Segmentation with Background-centric PromptingYuntian Bo, Yazhou Zhu, Piotr Koniusz, Haofeng ZhangCVPR 2026 · 1 citation
- Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale ApproachMir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. LittleCVPR 2024
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
