Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency
Yikai Wang, Chenjie Cao, Junqiu Yu, Ke Fan, Xiangyang Xue, Yanwei Fu
摘要
Recent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbitrary objects in the masked region that don't align with the rest of the image. (2) Color inconsistency: Inpainted regions often have color shifts that causes a smeared appearance, reducing image quality. Retraining the generative model could help solve these issues, but it's costly since state-of-the-art latent-based diffusion and rectified flow models require a three-stage training process: training a VAE, training a generative U-Net or transformer, and fine-tuning for inpainting. Instead, this paper proposes a post-processing approach, dubbed as ASUKA (Aligned Stable inpainting with UnKnown Areas prior), to improve inpainting models. To address unwanted object insertion, we leverage a Masked Auto-Encoder (MAE) for reconstruction-based priors. This mitigates object hallucination while maintaining the model's generation capabilities. To address color inconsistency, we propose a specialized VAE decoder that treats latent-to-image decoding as a local harmonization task, significantly reducing color shifts for color-consistent inpainting. We validate ASUKA on SD 1.5 and FLUX inpainting variants with Places2 and MISATO, our proposed diverse collection of datasets. Results show that ASUKA mitigates object hallucination and improves color consistency over standard diffusion and rectified flow models and other inpainting methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Follow-Your-Preference: Towards Preference-Aligned Image InpaintingYutao Shen, Junkun Yuan, Toru Aonishi, Hideki Nakayama 等ICLR 2026 · 被引用 21 次
- Precise Object and Effect Removal with Adaptive Target-Aware AttentionJixin Zhao, Zhouxia Wang, Peiqing Yang, Shangchen ZhouCVPR 2026 · 被引用 13 次
- Direct 3D-Aware Object Insertion via Decomposed Visual ProxiesJingbo Gong, Yikai Wang, Yushi Lan, Yuhao Wan 等ICML 2026 · 被引用 3 次
- You Only Erase Once: Erasing Anything without Bringing Unexpected ContentYixing Zhu, Qing Zhang, Wenju Xu, Wei-Shi ZhengCVPR 2026
- ReFocusEraser: Refocusing for Small Object Removal with Robust Context-Shadow RepairQingping Zheng, Bo Huang, Yang Liu, Haoyu Zhao 等ICLR 2026
它引用的顶会 Paper36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Diffusion Models as Masked AutoencodersChen Wei, Karttikeya Mangalam, Po-Yao Huang, Yanghao Li 等ICCV 2023 · 被引用 82 次
- SmartBrush: Text and Shape Guided Object Inpainting with Diffusion ModelShaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz 等CVPR 2023
- Your Latent Mask is Wrong: Pixel-Equivalent Latent Compositing for Diffusion ModelsRowan Bradbury, Dazhi ZhongCVPR 2026 · 被引用 5 次
- MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingJun Huang, Ting Liu, Yihang Wu, Xiaochao Qu 等CVPR 2025
- JPGNet: Joint Predictive Filtering and Generative Network for Image InpaintingQing Guo, Xiaoguang Li, Felix Juefei-Xu, Hongkai Yu 等ACM MM 2021 · 被引用 34 次
