Learning Contextual Transformer Network for Image Inpainting
Ye Deng, Siqi Hui, Sanping Zhou, Deyu Meng, Jinjun Wang
Abstract
Fully Convolutional Networks with attention modules have been proven effective for learning-based image inpainting. While many existing approaches could produce visually reasonable results, the generated images often show blurry textures or distorted structures around corrupted areas. The main reason is due to the fact that convolutional neural networks have limited capacity for modeling contextual information with long range dependencies. Although the attention mechanism can alleviate this problem to some extent, existing attention modules tend to emphasize similarities between the corrupted and the uncorrupted regions while ignoring the dependencies from within each of them. Hence, this paper proposes the Contextual Transformer Network (CTN) which not only learns relationships between the corrupted and the uncorrupted regions but also exploits their respective internal closeness. Besides, instead of a fully convolutional network, in our CTN, we stack several transformer blocks to replace convolution layers to better model the long range dependencies. Finally, by dividing the image into patches of different sizes, we propose a multi-scale multi-head attention module to better model the affinity among various image regions. Experiments on several benchmark datasets demonstrate superior performance by our proposed approach.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c355e18c-1552-497f-8e64-f88ac93eaed1Cited by top-tier papers4
- T-former: An Efficient Transformer for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng et al.ACM MM 2022 · 58 citations
- Contextual Outpainting with Object-Level Contrastive LearningJiacheng Li, Chang Chen, Zhiwei XiongCVPR 2022 · 10 citations
- View-consistent Object Removal in Radiance FieldsYiren Lu, Jing Ma, Yu YinACM MM 2024 · 3 citations
- Instruct2See: Learning to Remove Any Obstructions Across DistributionsJunhang Li, Yu Guo, Chuhua Xian, Shengfeng HeICML 2025
Related papers
- Atrous Pyramid Transformer with Spectral Convolution for Image InpaintingMuqi Huang, Lefei ZhangACM MM 2022 · 11 citations
- TransCNN-HAE: Transformer-CNN Hybrid AutoEncoder for Blind Image InpaintingHaoru Zhao, Zhaorui Gu, Bing Zheng, Haiyong ZhengACM MM 2022 · 30 citations
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
- Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional EncodingQiaole Dong, Chenjie Cao, Yanwei FuCVPR 2022 · 194 citations
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
