High-Fidelity Pluralistic Image Completion with Transformers
Ziyu Wan, Jingbo Zhang, Dongdong Chen, Jing Liao
Abstract
Image completion has made tremendous progress with convolutional neural networks (CNNs), because of their powerful texture modeling capacity. However, due to some inherent properties (e.g., local inductive prior, spatial-invariant kernels), CNNs do not perform well in understanding global structures or naturally support pluralistic completion. Recently, transformers demonstrate their power in modeling the long-term relationship and generating diverse results, but their computation complexity is quadratic to input length, thus hampering the application in processing high-resolution images. This paper brings the best of both worlds to pluralistic image completion: appearance prior reconstruction with transformer and texture replenishment with CNN. The former transformer recovers pluralistic coherent structures together with some coarse textures, while the latter CNN enhances the local texture details of coarse priors guided by the high-resolution masked images. The proposed method vastly outperforms state-of-the-art methods in terms of three aspects: 1) large performance boost on image fidelity even compared to deterministic completion methods; 2) better diversity and higher fidelity for pluralistic completion; 3) exceptional generalization ability on large masks and generic dataset, like ImageNet. Code and pre-trained models have been publicly released at https://github.com/raywzy/ICT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a77686dc-29c2-4e62-8c99-7693d5aa315cCited by top-tier papers58
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang et al.CVPR 2022 · 1,207 citations
- Improving Diffusion Models for Inverse Problems using Manifold ConstraintsHyungjin Chung, Byeongsu Sim, Dohoon Ryu, Jong Chul YeNeurIPS 2022 · 738 citations
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao et al.NeurIPS 2023 · 505 citations
- DiffIR: Efficient Diffusion Model for Image RestorationBin Xia, Yulun Zhang, Shiyin Wang, Yitong Wang et al.ICCV 2023 · 410 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Coherent Semantic Attention for Image InpaintingHongyu Liu, Bin Jiang, Yi Xiao, Chao YangICCV 2019 · 395 citations
- SC-FEGAN: Face Editing Generative Adversarial Network With User's Sketch and ColorYoungjoo Jo, Jongyoul ParkICCV 2019 · 325 citations
Related papers
- Incremental Transformer Structure Enhanced Image Inpainting with Masking Positional EncodingQiaole Dong, Chenjie Cao, Yanwei FuCVPR 2022 · 194 citations
- Delving Globally into Texture and Structure for Image InpaintingHaipeng Liu, Yang Wang, Meng Wang, Yong RuiACM MM 2022 · 26 citations
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan et al.ACM MM 2021 · 153 citations
- T-former: An Efficient Transformer for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng et al.ACM MM 2022 · 58 citations
- Large Occluded Human Image Completion via Image-Prior CooperatingHengrun Zhao, Yu Zeng, Huchuan Lu, Lijun WangAAAI 2024 · 2 citations
