Decoupled Textual Embeddings for Customized Image Generation
Yufei Cai, Yuxiang Wei, Zhilong Ji, Jinfeng Bai, Hu Han, Wangmeng Zuo
摘要
Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the subject-unrelated information (e.g., background and pose) with the learned concept, limiting the potential to compose concept into new scenes. To address these issues, we propose the DETEX, a novel approach that learns the disentangled concept embedding for flexible customized text-to-image generation. Unlike conventional methods that learn a single concept embedding from the given images, our DETEX represents each image using multiple word embeddings during training, i.e., a learnable image-shared subject embedding and several image-specific subject-unrelated embeddings. To decouple irrelevant attributes (i.e., background and pose) from the subject embedding, we further present several attribute mappers that encode each image as several image-specific subject-unrelated embeddings. To encourage these unrelated embeddings to capture the irrelevant information, we incorporate them with corresponding attribute words and propose a joint training strategy to facilitate the disentanglement. During inference, we only use the subject embedding for image generation, while selectively using image-specific embeddings to retain image-specified attributes. Extensive experiments demonstrate that the subject embedding obtained by our method can faithfully represent the target concept, while showing superior editability compared to the state-of-the-art methods. Our code will be available at https://github.com/PrototypeNx/DETEX.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Personalized Generation In Large Model Era: A SurveyYiyan Xu, Jinghao Zhang, Alireza Salemi, Xinting Hu 等ACL 2025 · 被引用 45 次
- PLACE: Adaptive Layout-Semantic Fusion for Semantic Image SynthesisZhengyao Lv, Yuxiang Wei, Wangmeng Zuo, Kwan-Yee K. WongCVPR 2024 · 被引用 14 次
- VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion ModelsYabo Zhang, Yuxiang Wei, Xianhui Lin, Zheng Hui 等AAAI 2025 · 被引用 3 次
- Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single ImageYuxuan Wang, Xuanyu Yi, Qingshan Xu, Yuan Zhou 等AAAI 2026 · 被引用 2 次
- Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image CustomizationLiyuan Ma, Xueji Fang, Guo-Jun QiACM MM 2024
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image GenerationHong Chen, Yipeng Zhang, Simin Wu, Xin Wang 等ICLR 2024 · 被引用 81 次
- Customizing Text-to-Image Generation with Inverted InteractionMengmeng Ge, Xu Jia, Takashi Isobe, Xiaomin Li 等ACM MM 2024 · 被引用 3 次
- ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token OptimizationMinseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon 等CVPR 2026 · 被引用 1 次
- Attention Calibration for Disentangled Text-to-Image PersonalizationYanbing Zhang, Mengping Yang, Qin Zhou, Zhe WangCVPR 2024 · 被引用 17 次
- TokenVerse: Versatile Multi-concept Personalization in Token Modulation SpaceDaniel Garibi, Shahar Yadin, Roni Paiss, Omer Tov 等SIGGRAPH 2025 · 被引用 12 次
