Pinco: Position-Induced Consistent Adapter for Diffusion Transformer in Foreground-Conditioned Inpainting
Guangben Lu, Yuzhen Du, Yizhe Tang, Zhimin Sun, Ran Yi, Yifan Qi, Tianyi Wang, Lizhuang Ma, Fangyuan Zou
Abstract
Foreground-conditioned inpainting aims to seamlessly fill the background region of an image by utilizing the provided foreground subject and a text description. While existing T2I-based image inpainting methods can be applied to this task, they suffer from issues of subject shape expansion, distortion, or impaired ability to align with the text description, resulting in inconsistencies between the visual elements and the text description. To address these challenges, we propose Pinco, a plug-and-play foreground-conditioned inpainting adapter that generates high-quality backgrounds with good text alignment while effectively preserving the shape of the foreground subject. Firstly, we design a Self-Consistent Adapter that integrates the foreground subject features into the layout-related self-attention layer, which helps to alleviate conflicts between the text and subject features by ensuring that the model can effectively consider the foreground subject's characteristics while processing the overall image layout. Secondly, we design a Decoupled Image Feature Extraction method that employs distinct architectures to extract semantic and spatial features separately, significantly improving subject feature extraction and ensuring high-quality preservation of the subject's shape. Thirdly, to ensure precise utilization of the extracted features and to focus attention on the subject region, we introduce a Shared Positional Embedding Anchor, greatly improving the model's understanding of subject features and boosting training efficiency. Extensive experiments demonstrate that our method achieves superior performance and efficiency in foreground-conditioned inpainting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8ddde56-632f-469e-9077-9b90dfece535Cited by top-tier papers3
- Social Debiasing for Fair Multi-Modal LLMsHarry Cheng, Yangyang Guo, Qing Guo, Ming-Hsuan Yang et al.ICCV 2025 · 1 citation
- FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow MatchingAndranik Sargsyan, Shant NavasardyanCVPR 2026
- ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background InpaintingYizhe Tang, Zhimin Sun, Yuzhen Du, Ran Yi et al.CVPR 2025
Builds on27
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterJianhui Zhang, Sheng Cheng, Qirui Sun, Jia Liu et al.ICCV 2025 · 1 citation
- Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image GeneratorChaehun Shin, Jooyoung Choi, Heeseung Kim, Sungroh YoonCVPR 2025
- MMFL: Multimodal Fusion Learning for Text-Guided Image InpaintingQing Lin, Bo Yan, Jichun Li, Weimin TanACM MM 2020 · 22 citations
- Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational AutoencoderWonwoong Cho, Yan-Ying Chen, Matthew Klenk, David I. Inouye et al.ICCV 2025
- IFAdapter: Instance Feature Control for Grounded Text-to-Image GenerationYinwei Wu, Xianpan Zhou, Bing Ma, Xuefeng Su et al.ICCV 2025 · 3 citations
