VRP-SAM: SAM with Visual Reference Prompt
Yanpeng Sun, Jiahui Chen, Shan Zhang, Xinyu Zhang, Qiang Chen, Gang Zhang, Errui Ding, Jingdong Wang, Zechao Li
Abstract
In this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Any-thing Model (SAM) to utilize annotated reference images as prompts for segmentation, creating the VRP-SAM model. In essence, VRP-SAM can utilize annotated reference images to comprehend specific objects and perform segmen-tation of specific objects in target image. It is note that the VRP encoder can support a variety of annotation for-mats for reference images, including point, box, scribble, and mask. VRP-SAM achieves a breakthrough within the SAM framework by extending its versatility and applicabil-ity while preserving SAM's inherent strengths, thus enhancing user-friendliness. To enhance the generalization abil-ity of VRP-SAM, the VRP encoder adopts a meta-learning strategy. To validate the effectiveness of VRP-SAM, we con-ducted extensive empirical studies on the Pascal and COCO datasets. Remarkably, VRP-SAM achieved state-of-the-art performance in visual reference segmentation with mini-mal learnable parameters. Furthermore, VRP-SAM demon-strates strong generalization capabilities, allowing it to per-form segmentation of unseen objects and enabling cross-domain segmentation. The source code and models will be available at https://github.com/syp2ysy/VRP-SAM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6ae6582-9bb9-4bfd-b247-0f7077a359c7Cited by top-tier papers26
- Bridge the Points: Graph-based Few-shot Segment Anything SemanticallyAnqi Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu et al.NeurIPS 2024 · 56 citations
- Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot SegmentationSuho Park, SuBeen Lee, Hyun Seok Seong, Jaejoon Yoo et al.AAAI 2025 · 9 citations
- SANSA: Unleashing the Hidden Semantics in SAM2 for Few-Shot SegmentationClaudia Cuttano, Gabriele Trivigno, Giuseppe Averta, Carlo MasoneNeurIPS 2025 · 9 citations
- RESAnything: Attribute Prompting for Arbitrary Referring SegmentationRuiqi Wang, Hao ZhangNeurIPS 2025 · 6 citations
- Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningJiawei Gu, Ziyue Qiao, Xinming Li, Zechao LiNeurIPS 2025 · 5 citations
Builds on18
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Feature Weighting and Boosting for Few-Shot SegmentationKhoi Nguyen, Sinisa TodorovicICCV 2019 · 402 citations
- Personalize Segment Anything Model with One ShotRenrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan et al.ICLR 2024 · 333 citations
Related papers
- MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationHang Xiong, Runmin Cong, Jinpeng Chen, Chen Zhang et al.ACM MM 2025
- ProSAM: Enhancing the Robustness of Sam-Based Visual Reference Segmentation with Probabilistic PromptsXiaoqi Wang, Clint Sebastian, Wenbin He, Liu RenICCV 2025 · 1 citation
- MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object SegmentationFu Rong, Meng Lan, Qian Zhang, Lefei ZhangICCV 2025 · 4 citations
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang et al.CVPR 2024 · 20 citations
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
