Mitigating Semantic Collapse in Generative Personalization with Test-Time Embedding Adjustment
Anh Bui, Thuy-Trang Vu, Trung Le, Junae Kim, Tamas Abraham, Rollin Omari, Amardeep Kaur, Dinh Q. Phung
摘要
In this paper, we investigate the semantic collapsing problem in generative personalization, an under-explored topic where the learned visual concept () gradually shifts from its original textual meaning and comes to dominate other concepts in multi-concept input prompts. This issue not only reduces the semantic richness of complex input prompts like"a photo of wearing glasses and playing guitar"into simpler, less contextually rich forms such as"a photo of "but also leads to simplified output images that fail to capture the intended concept. We identify the root cause as unconstrained optimisation, which allows the learned embedding to drift arbitrarily in the embedding space, both in direction and magnitude. To address this, we propose a simple yet effective training-free method that adjusts the magnitude and direction of pre-trained embedding at inference time, effectively mitigating the semantic collapsing problem. Our method is broadly applicable across different personalization methods and demonstrates significant improvements in text-image alignment in diverse use cases. Our code is anonymously published at https://github.com/tuananhbui89/Embedding-Adjustment
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image SynthesisShiyu Wu, Mingzhen Sun, Weining Wang, Yequan Wang 等ICLR 2026 · 被引用 7 次
- Encoder-based Domain Tuning for Fast Personalization of Text-to-Image ModelsRinon Gal, Moab Arar, Yuval Atzmon, Amit H. Bermano 等SIGGRAPH 2023 · 被引用 154 次
- AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image GenerationLianyu Pang, Jian Yin, Baoquan Zhao, Feize Wu 等NeurIPS 2024 · 被引用 18 次
- DreamMatcher: Appearance Matching Self-Attention for Semantically-Consistent Text-to-Image PersonalizationJisu Nam, Heesu Kim, DongJae Lee, Siyoon Jin 等CVPR 2024 · 被引用 21 次
- InstantBooth: Personalized Text-to-Image Generation without Test-Time FinetuningJing Shi, Wei Xiong, Zhe Lin, Hyun Joon JungCVPR 2024 · 被引用 115 次
