Referring Image Matting
Jizhizi Li, Jing Zhang, Dacheng Tao
Abstract
Different from conventional image matting, which either requires user-defined scribbles/trimap to extract a specific foreground object or directly extracts all the foreground objects in the image indiscriminately, we introduce a new task named Referring Image Matting (RIM) in this paper, which aims to extract the meticulous alpha matte of the specific object that best matches the given natural language description, thus enabling a more natural and simpler instruction for image matting. First, we establish a large-scale challenging dataset RefMatte by designing a comprehensive image composition and expression generation engine to automatically produce high-quality images along with diverse text attributes based on public datasets. RefMatte consists of 230 object categories, 47,500 images, 118,749 expression-region entities, and 474,996 expressions. Additionally, we construct a real-world test set with 100 high-resolution natural images and manually annotate complex phrases to evaluate the out-of-domain generalization abilities of RIM methods. Furthermore, we present a novel baseline method CLIPMat for RIM, including a context-embedded prompt, a text-driven semantic pop-up, and a multi-level details extractor. Extensive experiments on RefMatte in both keyword and expression settings validate the superiority of CLIPMat over representative methods. We hope this work could provide novel insights into image matting and encourage more followup studies. The dataset, code and models are available at https://github.com/JizhiziLi/RIM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang et al.MobiCom 2024 · 40 citations
- ParallelEdits: Efficient Multi-Aspect Text-Driven Image Editing with Attention GroupingMingzhen Huang, Jialing Cai, Shan Jia, Vishnu Suresh Lokhande et al.NeurIPS 2024 · 19 citations
- Serverless Federated AUPRC Optimization for Multi-Party Collaborative Imbalanced Data MiningXidong Wu, Zhengmian Hu, Jian Pei, Heng HuangKDD 2023 · 13 citations
- Unifying Automatic and Interactive Matting with Pretrained ViTsZixuan Ye, Wenze Liu, He Guo, Yujia Liang et al.CVPR 2024 · 7 citations
- VideoMaMa: Mask-Guided Video Matting via Generative PriorSangbeom Lim, Seoung Wug Oh, Gabriel Huang, Heeji Yoon et al.CVPR 2026 · 3 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or et al.ICCV 2021 · 1,437 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Semantic Image MattingYanan Sun, Chi-Keung Tang, Yu-Wing TaiCVPR 2021
- In-Context MattingHe Guo, Zixuan Ye, Zhiguo Cao, Hao LuCVPR 2024
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang et al.CVPR 2024 · 20 citations
- Matting by GenerationZhixiang Wang, Baiang Li, Jian Wang, Yu-Lun Liu et al.SIGGRAPH 2024 · 5 citations
- Referring Image Editing: Object-Level Image Editing via Referring ExpressionsChang Liu, Xiangtai Li, Henghui DingCVPR 2024
