Object Fusion via Diffusion Time-step for Customized Image Editing with Single Example
Xue Song, Zhongqi Yue, Jiequan Cui, Hanwang Zhang, Jingjing Chen
Abstract
We tackle the task of customized image editing using a text-conditioned Diffusion Model (DM). The goal is to fuse the subject in a reference image (e.g., sunglasses) with a source one (e.g., a boy), while retaining the fidelity of them both (e.g., the boy wearing the sunglasses). An intuitive approach, called LoRA fusion, first separately trains a DM LoRA for each image to encode its details. Then the two LoRAs are linearly combined by a weight to generate a fused image. Unfortunately, even through careful grid search or learning the weight, this approach still trades off the fidelity of one image against the other. We point out that the evil lies in the overlooked role of diffusion time-step in the generation process, i.e., a smaller time-step controls the generation of a more fine-grained attribute. For example, a large LoRA weight for the source may help preserve its fine-grained details (e.g., face attributes) at a small time-step, but could overpower the reference subject LoRA and lose the fidelity of its overall shape at a larger time-step. To address this deficiency, we propose TimeFusion, which learns a time-step-specific LoRA fusion weight that resolves the trade-off, i.e., generating the source and reference subject in high fidelity given their respective prompt. Then we can customize image editing using this weight and a target prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1e2be919-8181-499f-a60c-62fe6fe78703Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow TransformersYusuf Dalva, Hidir Yesiltepe, Pinar YanardagNeurIPS 2025 · 13 citations
- CustomTTT: Motion and Appearance Customized Video Generation via Test-Time TrainingXiuli Bi, Jian Lu, Bo Liu, Xiaodong Cun et al.AAAI 2025 · 10 citations
- T-LoRA: Single Image Diffusion Model Customization Without OverfittingVera Soboleva, Aibek Alanov, Andrey Kuznetsov, Konstantin SobolevAAAI 2026 · 9 citations
- CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free FusionYu Li, Yujun Cai, Chi ZhangCVPR 2026 · 2 citations
- Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image ModelsGihyun Kwon, Simon Jenni, Dingzeyu Li, Joon-Young Lee et al.CVPR 2024
