Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
Gihoon Kim, Hyungjin Park, Taesup Kim
摘要
Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the model's original generative capabilities. However, this process often leads to overfitting, where the model ignores the user's prompt and merely replicates the reference images. We attribute this issue to a fundamental misalignment between the true goals of personalization, which are subject fidelity and text alignment, and the training objectives of existing methods that fail to enforce both objectives simultaneously. Specifically, prior approaches often overlook the need to explicitly preserve the pretrained model's output distribution, resulting in distributional drift that undermines diversity and coherence. To resolve these challenges, we introduce a Lipschitz-based regularization objective that constrains parameter updates during personalization, ensuring bounded deviation from the original distribution. This promotes consistency with the pretrained model's behavior while enabling accurate adaptation to new concepts. Furthermore, our method offers a computationally efficient alternative to commonly used, resource-intensive sampling techniques. Through extensive experiments across diverse diffusion model architectures, we demonstrate that our approach achieves superior performance in both quantitative metrics and qualitative evaluations, consistently excelling in visual fidelity and prompt adherence. We further support these findings with comprehensive analyses, including ablation studies and visualizations. † Corresponding author Personalized Text-to-Image Generation. A central question of few-shot personalization has been how to adapt pretrained networks to new concepts with only a few subject-specific images. Textual Inversion (Gal et al., 2022; Voynov et al., 2023) encodes subject-specific information into learned text 1 This follows from the conditional distribution log p θ (x | c) = log p θ (x, c) -log p(c) = log ∫︁ p θ (x, z1:T , c) dz1:T -log p(c), where log p(c) is constant with respect to θ.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 被引用 13 次
- Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective DynamicsAnton Voronov, Mikhail Khoroshikh, Artem Babenko, Max RyabininNeurIPS 2023 · 被引用 8 次
- ClassDiffusion: More Aligned Personalization Tuning with Explicit Class GuidanceJiannan Huang, Jun Hao Liew, Hanshu Yan, Yuyang Yin 等ICLR 2025 · 被引用 1 次
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 等SIGGRAPH 2023 · 被引用 154 次
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu 等NeurIPS 2024 · 被引用 18 次
