SmartBrush: Text and Shape Guided Object Inpainting with Diffusion Model
Shaoan Xie, Zhifei Zhang, Zhe Lin, Tobias Hinz, Kun Zhang
摘要
Generic image inpainting aims to complete a corrupted image by borrowing surrounding information, which barely generates novel content. By contrast, multi-modal inpainting provides more flexible and useful controls on the inpainted content, e.g., a text prompt can be used to describe an object with richer attributes, and a mask can be used to constrain the shape of the inpainted object rather than being only considered as a missing area. We propose a new diffusion-based model named SmartBrush for completing a missing region with an object using both text and shape-guidance. While previous work such as DALLE-2 and Stable Diffusion can do text-guided inapinting they do not support shape guidance and tend to modify background texture surrounding the generated object. Our model incorporates both text and shape guidance with precision control. To preserve the background better, we propose a novel training and sampling strategy by augmenting the diffusion U-net with object-mask prediction. Lastly, we introduce a multi-task training strategy by jointly training inpainting with text-to-image generation to leverage more training data. We conduct extensive experiments showing that our model outperforms all baselines in terms of visual quality, mask controllability, and background preservation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper125
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang 等NeurIPS 2024 · 被引用 728 次
- Composer: Creative and Controllable Image Synthesis with Composable ConditionsLianghua Huang, Di Chen, Yu Liu, Yujun Shen 等ICML 2023 · 被引用 371 次
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi 等ICLR 2024 · 被引用 234 次
- FocalDreamer: Text-Driven 3D Editing via Focal-Fusion AssemblyYuhan Li, Yishun Dou, Yue Shi, Yu Lei 等AAAI 2024 · 被引用 91 次
- Zero-shot Image Editing with Reference ImitationXi Chen, Yutong Feng, Mengting Chen, Yiyang Wang 等NeurIPS 2024 · 被引用 80 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Uni-paint: A Unified Framework for Multimodal Image Inpainting with Pretrained Diffusion ModelShiyuan Yang, Xiaodong Chen, Jing LiaoACM MM 2023 · 被引用 65 次
- MagicPaint: Operate Anything for Image Inpainting with Diffusion ModelQinhong Yang, Dongdong Chen, Qi Chu, Tao Gong 等AAAI 2026
- Text-Guided Image InpaintingZijian Zhang, Zhou Zhao, Zhu Zhang, Baoxing Huai 等ACM MM 2020 · 被引用 17 次
- Brush2Prompt: Contextual Prompt Generator for Object InpaintingMang Tik Chiu, Yuqian Zhou, Lingzhi Zhang, Zhe Lin 等CVPR 2024
- SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout ControlJaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith 等CVPR 2024 · 被引用 7 次
