Plug-and-Play PPO: An Adaptive Point Prompt Optimizer Making SAM Greater
Xueyu Liu, Rui Wang, Yexin Lai, Guangze Shi, Feixue Shao, Fang Hao, Jianan Zhang, Jia Shen, Yongfei Wu, Wen Zheng
Abstract
Powered by extensive curated training data, the Segment Anything Model (SAM) demonstrates impressive generalization capabilities in open-world scenarios, effectively guided by user-provided prompts. However, the classagnostic characteristic of SAM renders its segmentation accuracy highly dependent on prompt quality. In this paper, we propose a novel plug-and-play dual-space Point Prompt Optimizer (PPO) designed to enhance prompt distribution through deep reinforcement learning (DRL)-based heterogeneous graph optimization. PPO optimizes initial prompts for any task without requiring additional training, thereby improving SAM's downstream segmentation performance. Specifically, PPO constructs a dual-space heterogeneous graph, leveraging the robust feature-matching capabilities of a foundational pre-trained model to create internal feature and physical distance matrices. A DRL policy network iteratively refines the distribution of prompt points, optimizing segmentation predictions. We conducted experiments on four public datasets. The ablation study explores the necessity and balance of optimizing prompts in both feature and physical spaces. The comparative study shows that PPO enables SAM to surpass recent one-shot methods. Additionally, experiments with different initial prompts demonstrate PPO's generality across prompts generated by various methods. In conclusion, PPO redefines the prompt optimization problem as a heterogeneous graph optimization task, using DRL to construct an effective, plugand-play prompt optimizer. This approach holds potential for broader applications across diverse segmentation tasks and provides a promising solution for point prompt optimization. The source code and demo are available at https://github.com/XueyuLiu/PPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59b7348a-e77d-4dfa-adda-d31c4af78d34Cited by top-tier papers2
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything ModelXueyu Liu, Xiaoyi Zhang, Meilin Liu, Guangze Shi et al.CVPR 2026 · 1 citation
- PromptPilot: Game-Theoretic Multi-Agent Prompt Optimization for Segment AnythingGuangze Shi, Yingjie Mi, Jia Shen, Feixue Shao et al.ICML 2026
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Personalize Segment Anything Model with One ShotRenrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan et al.ICLR 2024 · 333 citations
Related papers
- AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement LearningDuojun Huang, Xinyu Xiong, Jie Ma, Jichang Li et al.CVPR 2024
- AoP-SAM: Automation of Prompts for Efficient SegmentationYi Chen, Muyoung Son, Chuanbo Hua, Joo-Young KimAAAI 2025 · 9 citations
- VRP-SAM: SAM with Visual Reference PromptYanpeng Sun, Jiahui Chen, Shan Zhang, Xinyu Zhang et al.CVPR 2024 · 49 citations
- APSeg: Auto-Prompt Network for Cross-Domain Few-Shot Semantic SegmentationWeizhao He, Yang Zhang, Wei Zhuo, Linlin Shen et al.CVPR 2024
- MM-Prompt: Multi-modality and Multi-granularity Prompts for Few-Shot SegmentationHang Xiong, Runmin Cong, Jinpeng Chen, Chen Zhang et al.ACM MM 2025
