Continuously Masked Transformer for Image Inpainting
Keunsoo Ko, Chang-Su Kim
Abstract
A novel continuous-mask-aware transformer for image inpainting, called CMT, is proposed in this paper, which uses a continuous mask to represent the amounts of errors in tokens. First, we initialize a mask and use it during the self-attention. To facilitate the masked self-attention, we also introduce the notion of overlapping tokens. Second, we update the mask by modeling the error propagation during the masked self-attention. Through several masked self-attention and mask update (MSAU) layers, we predict initial inpainting results. Finally, we refine the initial results to reconstruct a more faithful image. Experimental results on multiple datasets show that the proposed CMT algorithm outperforms existing algorithms significantly. The source codes are available at https://github.com/keunsoo-ko/CMT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40b9e265-59af-4f63-a96b-6e86fb838453Cited by top-tier papers9
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li et al.CVPR 2026 · 14 citations
- BLS-GAN: A Deep Layer Separation Framework for Eliminating Bone Overlap in Conventional RadiographsHaolin Wang, Yafei Ou, Prasoon Ambalathankandy, Gen Ota et al.AAAI 2025 · 7 citations
- Perspective-Aware 3D Gaussian Inpainting with Multi-View ConsistencyYuxin Cheng, Binxiao Huang, Taiqiang Wu, Wenyong Zhou et al.ICCV 2025 · 1 citation
- Structure Matters: Tackling the Semantic Discrepancy in Diffusion Models for Image InpaintingHaipeng Liu, Yang Wang, Biao Qian, Meng Wang et al.CVPR 2024
- RORem: Training a Robust Object Remover with Human-in-the-LoopRuibin Li, Tao Yang, Song Guo, Lei ZhangCVPR 2025
Builds on13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
- Large Scale Image Completion via Co-Modulated Generative Adversarial NetworksShengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong et al.ICLR 2021 · 348 citations
Related papers
- Learning Contextual Transformer Network for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng et al.ACM MM 2021 · 29 citations
- Reduce Information Loss in Transformers for Pluralistic Image InpaintingQiankun Liu, Zhentao Tan, Dongdong Chen, Qi Chu et al.CVPR 2022 · 99 citations
- DLFormer: Discrete Latent Transformer for Video InpaintingJingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu et al.CVPR 2022 · 39 citations
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan et al.ACM MM 2021 · 153 citations
- TransCNN-HAE: Transformer-CNN Hybrid AutoEncoder for Blind Image InpaintingHaoru Zhao, Zhaorui Gu, Bing Zheng, Haiyong ZhengACM MM 2022 · 30 citations
