Spin: Diffusion-based Semantic Image Painting Through Independent Information Injection
Dantong Wu, Zhiqiang Chen, Tianjiao Du, Peipei Ran, Mengchao Bai, Kai Zhang
Abstract
Diffusion models have been utilized as powerful tools for various image editing tasks, including semantic image painting(SIP), which aims to generate content within masked regions conditioned on a reference image or text. SIP, especially those using images as condition, often suffers from three issues: semantic inconsistency, unnatural transitions and style inconsistency, which significantly hinder its practical application. To address these challenges, we propose a novel Semantic image Painting framework with INdependent INformation INjection(Spin). Specifically, we compute a saliency map to segregate the reference image into salient and non-salient components. We then filter out the non-salient information of it during the semantic embedding extraction phrase, and precisely inject the semantic embedding into the masked region instead of the whole image during the semantic generation phrase. Furthermore, we impose an additional style guidance to promote style consistency between background and foreground. Experimental results demonstrate that Spin achieve superior semantic similarity and image coherence across various styles, including realistic, pencil drawings, cartoon, and oil painting. Additionally, Spin offers diversity and editability, and can be integrated into other models that meet our prerequisites.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6e490fc9-6ca5-4592-b91c-e317c2f3984aBuilds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image SynthesisJunsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao et al.ICLR 2024 · 831 citations
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 339 citations
Related papers
- DiffEdit: Diffusion-based semantic image editing with mask guidanceGuillaume Couairon, Jakob Verbeek, Holger Schwenk, Matthieu CordICLR 2023 · 102 citations
- MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingJun Huang, Ting Liu, Yihang Wu, Xiaochao Qu et al.CVPR 2025
- Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive InjectionBowen Fu, Wei Wei, Jiaqi Tang, Jiangtao Nie et al.ICCV 2025 · 2 citations
- FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsYabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu et al.ICCV 2025 · 3 citations
- Brush2Prompt: Contextual Prompt Generator for Object InpaintingMang Tik Chiu, Yuqian Zhou, Lingzhi Zhang, Zhe Lin et al.CVPR 2024
