Adversarial Learning with Mask Reconstruction for Text-Guided Image Inpainting
Xingcai Wu, Yucheng Xie, Jiaqi Zeng, Zhenguo Yang, Yi Yu, Qing Li, Wenyin Liu
Abstract
Text-guided image inpainting aims to complete the corrupted patches coherent with both visual and textual context. On one hand, existing works focus on surrounding pixels of the corrupted patches without considering the objects in the image, resulting in the characteristics of objects described in text being painted on non-object regions. On the other hand, the redundant information in text may distract the generation of objects of interest in the restored image. In this paper, we propose an adversarial learning framework with mask reconstruction (ALMR) for image inpainting with textual guidance, which consists of a two-stage generator and dual discriminators. The two-stage generator aims to restore coarse-grained and fine-grained images, respectively. In particular, we devise a dual-attention module (DAM) to incorporate the word-level and sentence-level textual features as guidance on generating the coarse-grained and fine-grained details in the two stages. Furthermore, we design a mask reconstruction module (MRM) to penalize the restoration of the objects of interest with the given textual descriptions about the objects. For adversarial training, we exploit global and local discriminators for the whole image and corrupted patches, respectively. Extensive experiments conducted on CUB-200-2011, Oxford-102 and CelebA-HQ show the outperformance of the proposed ALMR (e.g., FID value is reduced from 29.69 to 14.69 compared with the state-of-the-art approach on CUB-200-2011). Codes are available at https://github.com/GaranWu/ALMR
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get d4d871a0-fd08-4091-a3c3-0a32e96cb6f6Cited by top-tier papers1
Ask how each one uses itRelated papers
- Text-Guided Neural Image InpaintingLisai Zhang, Qingcai Chen, Baotian Hu, Shuoran JiangACM MM 2020 · 53 citations
- MMFL: Multimodal Fusion Learning for Text-Guided Image InpaintingQing Lin, Bo Yan, Jichun Li, Weimin TanACM MM 2020 · 22 citations
- DAFT-GAN: Dual Affine Transformation Generative Adversarial Network for Text-Guided Image InpaintingJihoon Lee, Yunhong Min, Hwidong Kim, Sangtae AhnACM MM 2024 · 3 citations
- Text-Guided Image InpaintingZijian Zhang, Zhou Zhao, Zhu Zhang, Baoxing Huai et al.ACM MM 2020 · 17 citations
- Deep Multi-Resolution Mutual Learning for Image InpaintingHuan Zheng, Zhao Zhang, Haijun Zhang, Yi Yang et al.ACM MM 2022 · 14 citations
