SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation
Tianyuan Du, Haopeng Li, Zhen Fan, Jiarui Zhang, Panwang Pan, Yang Zhang
Abstract
Significant progress has been made in reasoning segmentation by combining multi-modal large language models (MLLMs) with the Segment Anything Model (SAM): the former excel in reasoning and vision-language alignment, while the latter offers powerful pixel-level understanding. However, current paradigms fall short in exploiting SAM's strengths, especially the ability to support iterative mask refinement by interactive segmentation, a process that human users can naturally perform. To bridge this gap, we introduce SAM-Veteran, an experienced mask-aware SAM agent capable of emulating human interaction with SAM via a reasoning-driven segmentation workflow that integrates (i) generating bounding boxes given image-query pairs for SAM input, (ii) proposing refinement points based on SAM-generated masks, and (iii) adaptively terminating the process. Aiming for this goal, we propose a multi-task reinforcement learning framework based on Group Relative Policy Optimization (GRPO), which enhances the MLLM's abilities in textual grounding and mask comprehension. Furthermore, we introduce a dynamic sampling strategy tailored for generating both boxes and points to stabilize training. Extensive experiments across diverse datasets show that SAM-Veteran achieves human-like interaction with SAM and establishes new state-of-the-art performance on both in-domain and out-of-domain benchmarks. Recent studies have investigated two primary MLLM-based paradigms for this task: (1) Supervised Fine-Tuning (SFT), where MLLMs generate special tokens that control a learnable segmentation head or decoder, thereby enabling end-to-end training as a unified model (Yan et al., 2024; Lai et al., 2024; Yan et al., 2025) ; and (2) Reinforcement Learning (RL), where MLLMs are optimized with reward signals for generating boxes and/or points that are then fed into Segment Anything Model (SAM) (Kirillov et al., 2023) to produce the final segmentation (Liu et al., 2025b; Huang et al., 2025) . While SFT-based methods effectively incorporate the reasoning capability of MLLMs into * Equal contribution. † Corresponding Author.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5fbdaa79-6246-4927-83b6-8c70982c9ea4Builds on36
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu et al.NeurIPS 2025 · 33 citations
- X-SAM: From Segment Anything to Any SegmentationHao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang et al.AAAI 2026 · 16 citations
- RSAgent: Learning to Reason and Act via Multi-Turn Tool Invocations for Text-Guided SegmentationXingqi He, Yujie Zhang, Shuyong Gao, Wenjie Li et al.ICML 2026 · 3 citations
- LENS: Learning to Segment Anything with Unified Reinforced ReasoningLianghui Zhu, Bin Ouyang, Yuxuan Zhang, Tianheng Cheng et al.AAAI 2026 · 7 citations
- MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionZhonghao Yan, Muxi Diao, Yuxuan Yang, Ruoyan Jing et al.AAAI 2026 · 4 citations
