MatteFormer: Transformer-Based Image Matting via Prior-Tokens
Gyutae Park, Sungjoon Son, Jaeyoung Yoo, Seho Kim, Nojun Kwak
Abstract
In this paper, we propose a transformer-based image matting model called MatteFormer, which takes full advantage of trimap information in the transformer block. Our method first introduces a prior-token which is a global representation of each trimap region (e.g. foreground, background and unknown). These prior-tokens are used as global priors and participate in the self-attention mechanism of each block. Each stage of the encoder is composed of PAST (Prior-Attentive Swin Transformer) block, which is based on the Swin Transformer block, but differs in a couple of aspects: 1) It has PA-WSA (Prior-Attentive Window Self-Attention) layer, performing self-attention not only with spatial-tokens but also with prior-tokens. 2) It has prior-memory which saves prior-tokens accumulatively from the previous blocks and transfers them to the next block. We evaluate our MatteFormer on the commonly used image matting datasets: Composition-1k and Distinctions-646. Experiment results show that our proposed method achieves state-of-the-art performance with a large margin. Our codes are available at https://github.com/ webtoon/matteformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72069841-1ce8-42cc-9a0a-5f44398d47c8Cited by top-tier papers21
- CLE Diffusion: Controllable Light Enhancement Diffusion ModelYuyang Yin, Dejia Xu, Chuangchuang Tan, Ping Liu et al.ACM MM 2023 · 77 citations
- X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionHanqing Zhao, Dianmo Sheng, Jianmin Bao, Dongdong Chen et al.ICML 2023 · 67 citations
- Neural-PBIR Reconstruction of Shape, Material, and IlluminationCheng Sun, Guangyan Cai, Zhengqin Li, Kai Yan et al.ICCV 2023 · 56 citations
- Unifying Automatic and Interactive Matting with Pretrained ViTsZixuan Ye, Wenze Liu, He Guo, Yujia Liang et al.CVPR 2024 · 7 citations
- dugMatting: Decomposed-Uncertainty-Guided MattingJiawei Wu, Changqing Zhang, Zuoyong Li, Huazhu Fu et al.ICML 2023 · 6 citations
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
Related papers
- SRFormer: Permuted Self-Attention for Single Image Super-ResolutionYupeng Zhou, Zhen Li, Chun-Le Guo, Song Bai et al.ICCV 2023
- Revisiting Context Aggregation for Image MattingQinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li et al.ICML 2024 · 6 citations
- Semantic Image MattingYanan Sun, Chi-Keung Tang, Yu-Wing TaiCVPR 2021
- Memory Efficient Matting with Adaptive Token RoutingYiheng Lin, Yihan Hu, Chenyi Zhang, Ting Liu et al.AAAI 2025 · 1 citation
- Disentangled Image MattingShaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Haibin Huang et al.ICCV 2019 · 127 citations
