Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE
Jialun Peng, Dong Liu, Songcen Xu, Houqiang Li
摘要
Given an incomplete image without additional constraint, image inpainting natively allows for multiple solutions as long as they appear plausible. Recently, multiplesolution inpainting methods have been proposed and shown the potential of generating diverse results. However, these methods have difficulty in ensuring the quality of each solution, e.g. they produce distorted structure and/or blurry texture. We propose a two-stage model for diverse inpainting, where the first stage generates multiple coarse results each of which has a different structure, and the second stage refines each coarse result separately by augmenting texture. The proposed model is inspired by the hierarchical vector quantized variational auto-encoder (VQ-VAE), whose hierarchical architecture disentangles structural and textural information. In addition, the vector quantization in VQ-VAE enables autoregressive modeling of the discrete distribution over the structural information. Sampling from the distribution can easily generate diverse and high-quality structures, making up the first stage of our model. In the second stage, we propose a structural attention module inside the texture generation network, where the module utilizes the structural information to capture distant correlations. We further reuse the VQ-VAE to calculate two feature losses, which help improve structure coherence and texture realism, respectively. Experimental results on CelebA-HQ, Places2, and ImageNet datasets show that our method not only enhances the diversity of the inpainting solutions but also improves the visual quality of the generated multiple images. Code and models are available at: https://github.com/USTC-JialunPeng/ Diverse-Structure-Inpainting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 等CVPR 2022 · 被引用 1,425 次
- Improving Diffusion Models for Inverse Problems using Manifold ConstraintsHyungjin Chung, Byeongsu Sim, Dohoon Ryu, Jong Chul YeNeurIPS 2022 · 被引用 738 次
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 等NeurIPS 2023 · 被引用 290 次
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan 等ACM MM 2021 · 被引用 153 次
- Bridging Global Context Interactions for High-Fidelity Image CompletionChuanxia Zheng, Tat-Jen Cham, Jianfei Cai, Dinh Q. PhungCVPR 2022 · 被引用 114 次
它引用的顶会 Paper6
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- Coherent Semantic Attention for Image InpaintingHongyu Liu, Bin Jiang, Yi Xiao, Chao YangICCV 2019 · 被引用 395 次
- StructureFlow: Image Inpainting via Structure-Aware Appearance FlowYurui Ren, Xiaoming Yu, Ruonan Zhang, Thomas H. Li 等ICCV 2019 · 被引用 356 次
- Contextual Residual Aggregation for Ultra High-Resolution Image InpaintingZili Yi, Qiang Tang, Shekoofeh Azizi, Daesik Jang 等CVPR 2020
- UCTGAN: Diverse Image Inpainting Based on Unsupervised Cross-Space TranslationLei Zhao, Qihang Mo, Sihuan Lin, Zhizhong Wang 等CVPR 2020
相关 Paper
- Locally Hierarchical Auto-Regressive Modeling for Image GenerationTackgeun You, Saehoon Kim, Chiheon Kim, Doyup Lee 等NeurIPS 2022 · 被引用 17 次
- Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision TransformersShiyue Cao, Yueqin Yin, Lianghua Huang, Yu Liu 等ICCV 2023 · 被引用 33 次
- Image Inpainting Based on Multi-frequency Probabilistic Inference ModelJin Wang, Chen Wang, Qingming Huang, Yunhui Shi 等ACM MM 2020 · 被引用 7 次
- Towards Accurate Image Coding: Improved Autoregressive Image Generation with Dynamic Vector QuantizationMengqi Huang, Zhendong Mao, Zhuowei Chen, Yongdong ZhangCVPR 2023
- Hierarchical Sketch Induction for Paraphrase GenerationTom Hosking, Hao Tang, Mirella LapataACL 2022
