AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
Lianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu, Fu Lee Wang, Qing Li, Xudong Mao
摘要
Recent advances in text-to-image models have enabled high-quality personalized image synthesis of user-provided concepts with flexible textual control. In this work, we analyze the limitations of two primary techniques in text-to-image personalization: Textual Inversion and DreamBooth. When integrating the learned concept into new prompts, Textual Inversion tends to overfit the concept, while DreamBooth often overlooks it. We attribute these issues to the incorrect learning of the embedding alignment for the concept. We introduce AttnDreamBooth, a novel approach that addresses these issues by separately learning the embedding alignment, the attention map, and the subject identity in different training stages. We also introduce a cross-attention map regularization term to enhance the learning of the attention map. Our method demonstrates significant improvements in identity preservation and text alignment compared to the baseline methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- ImageRAG: Dynamic Image Retrieval for Reference-Guided Image GenerationRotem Shalev-Arkushin, Rinon Gal, Amit Bermano, Ohad FriedICLR 2026 · 被引用 25 次
- CoRe: Context-Regularized Text Embedding Learning for Text-to-Image PersonalizationFeize Wu, Yun Pang, Junyi Zhang, Lianyu Pang 等AAAI 2025 · 被引用 13 次
- T-LoRA: Single Image Diffusion Model Customization Without OverfittingVera Soboleva, Aibek Alanov, Andrey Kuznetsov, Konstantin SobolevAAAI 2026 · 被引用 9 次
- POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image GenerationEvans Xu Han, Alice Qian Zhang, Haiyi Zhu, Hong Shen 等UIST 2025 · 被引用 5 次
它引用的顶会 Paper43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- DreamBooth3D: Subject-Driven Text-to-3D GenerationAmit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer 等ICCV 2023 · 被引用 280 次
- Nested Attention: Semantic-aware Attention Values for Concept PersonalizationOr Patashnik, Rinon Gal, Daniil Ostashev, Sergey Tulyakov 等SIGGRAPH 2025 · 被引用 6 次
- DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationJisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin 等CVPR 2024 · 被引用 21 次
- Is This Loss Informative? Faster Text-to-Image Customization by Tracking Objective DynamicsAnton Voronov, Mikhail Khoroshikh, Artem Babenko, Max RyabininNeurIPS 2023 · 被引用 8 次
- Compositional Inversion for Stable Diffusion ModelsXulu Zhang, Xiao-Yong Wei, Jinlin Wu, Tianyi Zhang 等AAAI 2024 · 被引用 26 次
