Reduce Information Loss in Transformers for Pluralistic Image Inpainting
Qiankun Liu, Zhentao Tan, Dongdong Chen, Qi Chu, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan, Nenghai Yu
摘要
Transformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer from information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration, incurring information loss and extra misalignment for the boundaries of masked regions. 2) They quantize 256 3 RGB pixels to a small number (such as 512) of quantized pixels. The indices of quantized pixels are used as tokens for the inputs and prediction targets of transformer. Although an extra CNN network is used to upsample and refine the low-resolution results, it is difficult to retrieve the lost information back. To keep input information as much as possible, we propose a new transformer based framework "PUT". Specifically, to avoid input downsampling while maintaining the computation efficiency, we design a patch-based auto-encoder P-VQVAE, where the encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by quantization, an Un-Quantized Transformer (UQ-Transformer) is applied, which directly takes the features from P-VQVAE encoder as input without quantization and regards the quantized tokens only as prediction targets. Extensive experiments show that PUT greatly outperforms state-of-the-art methods on image fidelity, especially for large masked regions and complex large-scale datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsShihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao 等NeurIPS 2023 · 被引用 505 次
- Continuously Masked Transformer for Image InpaintingKeunsoo Ko, Chang-Su KimICCV 2023 · 被引用 54 次
- Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive LearningZhiwen Zuo, Lei Zhao, Ailin Li, Zhizhong Wang 等AAAI 2023 · 被引用 36 次
- BVINet: Unlocking Blind Video Inpainting With Zero AnnotationsZhiliang Wu, Kerui Chen, Kun Li, Hehe Fan 等ICCV 2025 · 被引用 30 次
- Learning Image-Adaptive Codebooks for Class-Agnostic Image RestorationKechun Liu, Yitong Jiang, Inchang Choi, Jinwei GuICCV 2023 · 被引用 24 次
它引用的顶会 Paper20
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped WindowsXiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang 等CVPR 2022 · 被引用 1,207 次
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu 等CVPR 2022 · 被引用 600 次
相关 Paper
- Don't Look into the Dark: Latent Codes for Pluralistic Image InpaintingHaiwei Chen, Yajie ZhaoCVPR 2024
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi 等CVPR 2022 · 被引用 382 次
- Delving Globally into Texture and Structure for Image InpaintingHaipeng Liu, Yang Wang, Meng Wang, Yong RuiACM MM 2022 · 被引用 26 次
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu 等AAAI 2024 · 被引用 85 次
- Diverse Image Inpainting with Bidirectional and Autoregressive TransformersYingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan 等ACM MM 2021 · 被引用 153 次
