IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation
Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian L. Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Daniel G. Aliaga
Abstract
Generative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality. Project page: https://song630.github.io/IMPRINT-Project-Page/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f41fc062-02f8-4bfb-81bf-fbfc6afed3edCited by top-tier papers41
- Zero-shot Image Editing with Reference ImitationXi Chen, Yutong Feng, Mengting Chen, Yiyang Wang et al.NeurIPS 2024 · 80 citations
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu et al.ACL 2025 · 45 citations
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang et al.ICLR 2026 · 34 citations
- Training-Free Industrial Defect Generation with Diffusion ModelsRuyi Xu, Yen-Tzu Chiu, Tai-I Chen, Oscar Chew et al.ICCV 2025 · 8 citations
- Bifröst: 3D-Aware Image Compositing with Language InstructionsLingxiao Li, Kaixiong Gong, Wei-Hong Li, Xili Dai et al.NeurIPS 2024 · 5 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Zero-Shot Depth Aware Image Editing With Diffusion ModelsRishubh Parihar, Sachidanand VS, R. Venkatesh BabuICCV 2025 · 3 citations
- Direct 3D-Aware Object Insertion via Decomposed Visual ProxiesJingbo Gong, Yikai Wang, Yushi Lan, Yuhao Wan et al.ICML 2026 · 3 citations
- Omnipaint: Mastering Object-Oriented Editing Via Disentangled Insertion-Removal InpaintingYongsheng Yu, Ziyun Zeng, Haitian Zheng, Jiebo LuoICCV 2025 · 3 citations
- Teleportraits: Training-Free People Insertion Into Any SceneJialu Gao, K. J. Joseph, Fernando De la TorreICCV 2025
- A Diffusion-Based Framework for Occluded Object MovementZheng-Peng Duan, Jiawei Zhang, Siyu Liu, Zheng Lin et al.AAAI 2025 · 7 citations
