Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE
Jialun Peng, Dong Liu, Songcen Xu, Houqiang Li
Abstract
Given an incomplete image without additional constraint, image inpainting natively allows for multiple solutions as long as they appear plausible. Recently, multiplesolution inpainting methods have been proposed and shown the potential of generating diverse results. However, these methods have difficulty in ensuring the quality of each solution, e.g. they produce distorted structure and/or blurry texture. We propose a two-stage model for diverse inpainting, where the first stage generates multiple coarse results each of which has a different structure, and the second stage refines each coarse result separately by augmenting texture. The proposed model is inspired by the hierarchical vector quantized variational auto-encoder (VQ-VAE), whose hierarchical architecture disentangles structural and textural information. In addition, the vector quantization in VQ-VAE enables autoregressive modeling of the discrete distribution over the structural information. Sampling from the distribution can easily generate diverse and high-quality structures, making up the first stage of our model. In the second stage, we propose a structural attention module inside the texture generation network, where the module utilizes the structural information to capture distant correlations. We further reuse the VQ-VAE to calculate two feature losses, which help improve structure coherence and texture realism, respectively. Experimental results on CelebA-HQ, Places2, and ImageNet datasets show that our method not only enhances the diversity of the inpainting solutions but also improves the visual quality of the generated multiple images. Code and models are available at: https://github.com/USTC-JialunPeng/ Diverse-Structure-Inpainting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7570b24-8be7-4537-bb15-5b189c3cf1baCited by top-tier papers51
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- Improving Diffusion Models for Inverse Problems using Manifold ConstraintsHyungjin Chung, Byeongsu Sim, Dohoon Ryu, Jong Chul YeNeurIPS 2022 · 738 citations
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan et al.ACM MM 2021 · 153 citations
- Bridging Global Context Interactions for High-Fidelity Image CompletionChuanxia Zheng, Tat-Jen Cham, Jianfei Cai, Dinh Q. PhungCVPR 2022 · 114 citations
Builds on6
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Coherent Semantic Attention for Image InpaintingHongyu Liu, Bin Jiang, Yi Xiao, Chao YangICCV 2019 · 395 citations
- StructureFlow: Image Inpainting via Structure-Aware Appearance FlowYurui Ren, Xiaoming Yu, Ruonan Zhang, Thomas H. Li et al.ICCV 2019 · 356 citations
- Contextual Residual Aggregation for Ultra High-Resolution Image InpaintingZili Yi, Qiang Tang, Shekoofeh Azizi, Daesik Jang et al.CVPR 2020
- UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationLei Zhao, Qihang Mo, Sihuan Lin, Zhizhong Wang et al.CVPR 2020
Related papers
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee et al.NeurIPS 2022 · 17 citations
- Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision TransformersShiyue Cao, Yueqin Yin, Lianghua Huang, Yu Liu et al.ICCV 2023 · 33 citations
- Image Inpainting Based on Multi-frequency Probabilistic Inference ModelJin Wang, Chen Wang, Qingming Huang, Yunhui Shi et al.ACM MM 2020 · 7 citations
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Hierarchical Sketch Induction for Paraphrase GenerationTom Hosking, Hao Tang, Mirella LapataACL 2022
