MatteFormer: Transformer-Based Image Matting via Prior-Tokens
Gyutae Park, Sungjoon Son, Jaeyoung Yoo, Seho Kim, Nojun Kwak
摘要
In this paper, we propose a transformer-based image matting model called MatteFormer, which takes full advantage of trimap information in the transformer block. Our method first introduces a prior-token which is a global representation of each trimap region (e.g. foreground, background and unknown). These prior-tokens are used as global priors and participate in the self-attention mechanism of each block. Each stage of the encoder is composed of PAST (Prior-Attentive Swin Transformer) block, which is based on the Swin Transformer block, but differs in a couple of aspects: 1) It has PA-WSA (Prior-Attentive Window Self-Attention) layer, performing self-attention not only with spatial-tokens but also with prior-tokens. 2) It has prior-memory which saves prior-tokens accumulatively from the previous blocks and transfers them to the next block. We evaluate our MatteFormer on the commonly used image matting datasets: Composition-1k and Distinctions-646. Experiment results show that our proposed method achieves state-of-the-art performance with a large margin. Our codes are available at https://github.com/ webtoon/matteformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- CLE Diffusion: Controllable Light Enhancement Diffusion ModelYuyang Yin, Dejia Xu, Chuangchuang Tan, Ping Liu 等ACM MM 2023 · 被引用 77 次
- X-Paste: Revisiting Scalable Copy-Paste for Instance Segmentation using CLIP and StableDiffusionHanqing Zhao, Dianmo Sheng, Jianmin Bao, Dongdong Chen 等ICML 2023 · 被引用 67 次
- Neural-PBIR Reconstruction of Shape, Material, and IlluminationCheng Sun, Guangyan Cai, Zhengqin Li, Kai Yan 等ICCV 2023 · 被引用 56 次
- Unifying Automatic and Interactive Matting with Pretrained ViTsZixuan Ye, Wenze Liu, He Guo, Yujia Liang 等CVPR 2024 · 被引用 7 次
- dugMatting: Decomposed-Uncertainty-Guided MattingJiawei Wu, Changqing Zhang, Zuoyong Li, Huazhu Fu 等ICML 2023 · 被引用 6 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- SRFormer: Permuted Self-Attention for Single Image Super-ResolutionYupeng Zhou, Zhen Li, Chun-Le Guo, Song Bai 等ICCV 2023
- Revisiting Context Aggregation for Image MattingQinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li 等ICML 2024 · 被引用 6 次
- Semantic Image MattingYanan Sun, Chi-Keung Tang, Yu-Wing TaiCVPR 2021
- Memory Efficient Matting with Adaptive Token RoutingYiheng Lin, Yihan Hu, Chenyi Zhang, Ting Liu 等AAAI 2025 · 被引用 1 次
- Disentangled Image MattingShaofan Cai, Xiaoshuai Zhang, Haoqiang Fan, Haibin Huang 等ICCV 2019 · 被引用 127 次
