Spin: Diffusion-based Semantic Image Painting Through Independent Information Injection
Dantong Wu, Zhiqiang Chen, Tianjiao Du, Peipei Ran, Mengchao Bai, Kai Zhang
摘要
Diffusion models have been utilized as powerful tools for various image editing tasks, including semantic image painting(SIP), which aims to generate content within masked regions conditioned on a reference image or text. SIP, especially those using images as condition, often suffers from three issues: semantic inconsistency, unnatural transitions and style inconsistency, which significantly hinder its practical application. To address these challenges, we propose a novel Semantic image Painting framework with INdependent INformation INjection(Spin). Specifically, we compute a saliency map to segregate the reference image into salient and non-salient components. We then filter out the non-salient information of it during the semantic embedding extraction phrase, and precisely inject the semantic embedding into the masked region instead of the whole image during the semantic generation phrase. Furthermore, we impose an additional style guidance to promote style consistency between background and foreground. Experimental results demonstrate that Spin achieve superior semantic similarity and image coherence across various styles, including realistic, pencil drawings, cartoon, and oil painting. Additionally, Spin offers diversity and editability, and can be integrated into other models that meet our prerequisites.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 等CVPR 2022 · 被引用 1,425 次
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao 等ICLR 2024 · 被引用 831 次
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 被引用 670 次
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 被引用 339 次
相关 Paper
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 被引用 102 次
- MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingJun Huang, Ting Liu, Yihang Wu, Xiaochao Qu 等CVPR 2025
- Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionBowen Fu, Wei Wei, Jiaqi Tang, Jiangtao Nie 等ICCV 2025 · 被引用 2 次
- FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsYabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu 等ICCV 2025 · 被引用 3 次
- Brush2Prompt: Contextual Prompt Generator for Object InpaintingMang Tik Chiu, Yuqian Zhou, Lingzhi Zhang, Zhe Lin 等CVPR 2024
