Unifying Automatic and Interactive Matting with Pretrained ViTs
Zixuan Ye, Wenze Liu, He Guo, Yujia Liang, Chaoyi Hong, Hao Lu, Zhiguo Cao
摘要
Automatic and interactive matting largely improve image matting by respectively alleviating the need for auxil-iary input and enabling object selection. Due to different settings on whether prompts exist, they either suffer from weakness in instance completeness or region details. Also, when dealing with different scenarios, directly switching between the two matting models introduces inconvenience and higher workload. Therefore, we wonder whether we can al-leviate the limitations of both settings while achieving unification to facilitate more convenient use. Our key idea is to offer saliency guidance for automatic mode to enable its attention to detailed regions, and also refine the instance completeness in interactive mode by replacing the binary mask guidance with a more probabilistic form. With different guidance for each mode, we can achieve unification through adaptable guidance, defined as saliency information in automatic mode and user cue for interactive one. It is instantiated as candidate feature in our method, an automatic switch for class token in pretrained ViTs and average feature of user prompts, controlled by the existence of user prompts. Then we use the candidate feature to generate a probabilistic similarity map as the guidance to alleviate the over-reliance on binary mask. Extensive experiments show that our method can adapt well to both automatic and inter-active scenarios with more light-weight framework. Code available at github.com/coconut/SMat.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Edit2Perceive: Image Editing Diffusion Models Are Strong Dense PerceiversYiqing Shi, Yiren Song, Mike Zheng ShouCVPR 2026 · 被引用 4 次
- ZIM: Zero-Shot Image Matting for AnythingBeomyoung Kim, Chanyong Shin, Joonhyun Jeong, Hyungsik Jung 等ICCV 2025 · 被引用 2 次
- SDMATTE: Grafting Diffusion Models for Interactive MattingLongfei Huang, Yu Liang, Hao Zhang, Jinwei Chen 等ICCV 2025 · 被引用 1 次
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang 等ICLR 2026
- Segment and Matte Anything in a Unified ModelZezhong Fan, Xiaohan Li, Topojoy Biswas, Kaushiki Nag 等AAAI 2026
它引用的顶会 Paper18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- MODNet: Real-Time Trimap-Free Portrait Matting via Objective DecompositionZhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan 等AAAI 2022 · 被引用 220 次
- Indices Matter: Learning to Index for Deep Image MattingHao Lu, Yutong Dai, Chunhua Shen, Songcen XuICCV 2019 · 被引用 206 次
- Natural Image Matting via Guided Contextual AttentionYaoyi Li, Hongtao LuAAAI 2020 · 被引用 189 次
相关 Paper
- Situational Perception Guided Image MattingBo Xu, Jiake Xie, Han Huang, Ziwen Li 等ACM MM 2022 · 被引用 4 次
- In-Context MattingHe Guo, Zixuan Ye, Zhiguo Cao, Hao LuCVPR 2024
- MaGGIe: Masked Guided Gradual Human Instance MattingChuong Huynh, Seoung Wug Oh, Abhinav Shrivastava, Joon-Young LeeCVPR 2024
- Virtual Multi-Modality Self-Supervised Foreground Matting for Human-Object InteractionBo Xu, Han Huang, Cheng Lu, Ziwen Li 等ICCV 2021 · 被引用 7 次
- Mask-Guided Matting in the WildKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon 等CVPR 2023
