Segment and Matte Anything in a Unified Model
Zezhong Fan, Xiaohan Li, Topojoy Biswas, Kaushiki Nag, Kannan Achan
Abstract
Segment Anything (SAM) has recently pushed the boundaries of segmentation by demonstrating remarkable zero-shot generalization and flexible prompting after training on over one billion masks. Despite this, its mask prediction accuracy often falls short of the precision required in real-world applications. While several refinement modules have been proposed to boost SAM’s segmentation quality, achieving highly accurate object delineation within a single, unified framework remains an open challenge. Furthermore, interactive image matting—which aims to generate fine-grained alpha mattes guided by diverse user hints—has not yet been explored in the context of SAM. Insights from recent studies highlight strong correlations between segmentation and matting, suggesting the feasibility of a unified model capable of both tasks.
In this paper, we introduce Segment And Matte Anything (SAMA), a lightweight extension of SAM that delivers high-quality interactive image segmentation and matting with minimal extra parameters or computational cost. Our Multi-View Localization Encoder (MVLE) captures detailed features from local views, while the Localization Adapter (Local-Adapter) refines mask outputs by recovering subtle boundary details. We also incorporate two prediction heads for each task into the architecture to generate segmentation and matting tasks, simultaneously. Trained on a diverse dataset aggregated from publicly available sources, SAMA achieves state-of-the-art performance across multiple segmentation and matting benchmarks, showcasing its adaptability and effectiveness in a wide range of downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 817e2495-5dcd-47d3-9e0c-7042ba65e24cBuilds on30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
- Segment Anything in High QualityLei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu et al.NeurIPS 2023 · 709 citations
- Towards High-Resolution Salient Object DetectionYi Zeng, Pingping Zhang, Zhe Lin, Jianming Zhang et al.ICCV 2019 · 232 citations
- MODNet: Real-Time Trimap-Free Portrait Matting via Objective DecompositionZhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan et al.AAAI 2022 · 220 citations
Related papers
- Towards Fine-Grained Interactive Segmentation in Images and VideosYuan Yao, Qiushi Yang, Miaomiao Cui, Liefeng BoICCV 2025 · 2 citations
- SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything ModelChongkai Yu, Ting Liu, Anqi Li, Xiaochao Qu et al.CVPR 2025
- ZIM: Zero-Shot Image Matting for AnythingBeomyoung Kim, Chanyong Shin, Joonhyun Jeong, Hyungsik Jung et al.ICCV 2025 · 2 citations
- Segment Anything with Precise InteractionMengzhen Liu, Mengyu Wang, Henghui Ding, Yilong Xu et al.ACM MM 2024 · 2 citations
- AoP-SAM: Automation of Prompts for Efficient SegmentationYi Chen, Muyoung Son, Chuanbo Hua, Joo-Young KimAAAI 2025 · 9 citations
