Stable Segment Anything Model
Qi Fan, Xin Tao, Lei Ke, Mingqiao Ye, Di Zhang, Pengfei Wan, Yu-Wing Tai, Chi-Keung Tang
Abstract
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM's segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM's mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM's mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner, facilitated by our effective robust training strategy (RTS). During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM's segmentation stability across a wide range of prompt qualities, while 2) retaining SAM's powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation (by 1 training epoch). Extensive experiments across multiple datasets validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes will be released upon acceptance. https://github.com/fanq15/Stable-SAM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07580a14-afaa-4419-bfd5-ff3f46d0fe85Cited by top-tier papers3
- Q-Norm: Robust Representation Learning via Quality-Adaptive NormalizationLanning Zhang, Ying Zhou, Fei Gao, Ziyun Li et al.ICCV 2025 · 1 citation
- 3DTeethSAM: Taming SAM2 for 3D Teeth SegmentationZhiguo Lu, Jianwen Lou, Mingjun Ma, Hairong Jin et al.AAAI 2026 · 1 citation
- QuARF: Quality-Adaptive Receptive Fields for Degraded Image PerceptionFei Gao, Ying Zhou, Ziyun Li, Wenwang Han et al.AAAI 2025
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Vision Transformer with Deformable AttentionZhuofan Xia, Xuran Pan, Shiji Song, Li Erran Li et al.CVPR 2022 · 835 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
Related papers
- RobustSAM: Segment Anything Robustly on Degraded ImagesWei-Ting Chen, Yu-Jiet Vong, Sy-Yen Kuo, Sizhuo Ma et al.CVPR 2024
- FocSAM: Delving Deeply into Focused Objects in Segmenting AnythingYou Huang, Zongyu Lan, Liujuan Cao, Xianming Lin et al.CVPR 2024
- Attack for Defense: Adversarial Agents for Point Prompt Optimization Empowering Segment Anything ModelXueyu Liu, Xiaoyi Zhang, Meilin Liu, Guangze Shi et al.CVPR 2026 · 1 citation
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Unleashing the Potential of SAM for Medical Adaptation via Hierarchical DecodingZhiheng Cheng, Qingyue Wei, Hongru Zhu, Yan Wang et al.CVPR 2024
