Personalized Residuals for Concept-Driven Text-to-Image Generation
Cusuh Ham, Matthew Fisher, James Hays, Nicholas I. Kolkin, Yuchen Liu, Richard Zhang, Tobias Hinz
摘要
We present personalized residuals and localized attention-guided sampling for efficient concept-driven generation using text-to-image diffusion models. Our method first represents concepts by freezing the weights of a pretrained text-conditioned diffusion model and learning low-rank residuals for a small subset of the model's layers. The residual-based approach then directly enables application of our proposed sampling technique, which applies the learned residuals only in areas where the concept is localized via cross-attention and applies the original diffusion weights in all other regions. Localized sampling therefore combines the learned identity of the concept with the existing generative prior of the underlying diffusion model. We show that personalized residuals effectively capture the identity of a concept in ∼3 minutes on a single GPU without the use of regularization images and with fewer parameters than previous models, and localized sampling allows using the original model as strong prior for large parts of the image.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Fine-Tuning Visual Autoregressive Models for Subject-Driven GenerationJiwoo Chung, Sangeek Hyun, Hyunjun Kim, Eunseo Koh 等ICCV 2025 · 被引用 1 次
- Improving Personalized Search with Regularized Low-Rank Parameter UpdatesFiona Ryan, Josef Sivic, Fabian Caba Heilbron, Judy Hoffman 等CVPR 2025
- ID-Sim: An Identity-Focused Similarity MetricJulia Chae, Nick Kolkin, Jui-Hsien Wang, Richard Zhang 等CVPR 2026
- Multi-subject Open-set Personalization in Video GenerationTsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang 等CVPR 2025
- RAP: Retrieval-Augmented Personalization for Multimodal Large Language ModelsHaoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li 等CVPR 2025
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Steering Guidance for Personalized Text-to-Image Diffusion ModelsSunghyun Park, Seokeon Choi, Hyoungwoo Park, Sungrack YunICCV 2025 · 被引用 2 次
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 等SIGGRAPH 2023 · 被引用 154 次
- UniversalBooth: Model-Agnostic Personalized Text-To-Image GenerationSonghua Liu, Ruonan Yu, Xinchao WangICCV 2025 · 被引用 2 次
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman 等ICCV 2023 · 被引用 327 次
- Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional DriftGihoon Kim, Hyungjin Park, Taesup KimICLR 2026 · 被引用 1 次
