DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
Hong Chen, Yipeng Zhang, Simin Wu, Xin Wang, Xuguang Duan, Yuwei Zhou, Wenwu Zhu
摘要
Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to the failure of subject-driven text-to-image generation as follows: (i) the identity-irrelevant information hidden in the entangled embedding may dominate the generation process, resulting in the generated images heavily dependent on the irrelevant information while ignoring the given text descriptions; (ii) the identityrelevant information carried in the entangled embedding can not be appropriately preserved, resulting in identity change of the subject in the generated images. To tackle the problems, we propose DisenBooth, an identity-preserving disentangled tuning framework for subject-driven text-to-image generation. Specifically, Dis-enBooth finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise each image, DisenBooth instead utilizes disentangled embeddings to respectively preserve the subject identity and capture the identity-irrelevant information. We further design the novel weak denoising and contrastive embedding auxiliary tuning objectives to achieve the disentanglement. Extensive experiments show that our proposed DisenBooth framework outperforms baseline models for subject-driven text-to-image generation with the identity-preserved embedding. Additionally, by combining the identity-preserved embedding and identity-irrelevant embedding, DisenBooth demonstrates more generation flexibility and controllability 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Subject-Diffusion: Open Domain Personalized Text-to-Image Generation without Test-time Fine-tuningJian Ma, Junhao Liang, Chen Chen, Haonan LuSIGGRAPH 2024 · 被引用 71 次
- MultiBooth: Towards Generating All Your Concepts in an Image from TextChenyang Zhu, Kai Li, Yue Ma, Chunming He 等AAAI 2025 · 被引用 52 次
- SSR-Encoder: Encoding Selective Subject Representation for Subject-Driven GenerationYuxuan Zhang, Yiren Song, Jiaming Liu, Rui Wang 等CVPR 2024 · 被引用 34 次
- LLM4DyG: Can Large Language Models Solve Spatial-Temporal Problems on Dynamic Graphs?Zeyang Zhang, Xin Wang, Ziwei Zhang, Haoyang Li 等KDD 2024 · 被引用 32 次
- Curriculum Co-disentangled Representation Learning across Multiple Environments for Social RecommendationXin Wang, Zirui Pan, Yuwei Zhou, Hong Chen 等ICML 2023 · 被引用 31 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- DisenStudio: Customized Multi-Subject Text-to-Video Generation with Disentangled Spatial ControlHong Chen, Xin Wang, Yipeng Zhang, Yuwei Zhou 等ACM MM 2024 · 被引用 10 次
- Decoupled Textual Embeddings for Customized Image GenerationYufei Cai, Yuxiang Wei, Zhilong Ji, Jinfeng Bai 等AAAI 2024 · 被引用 24 次
- PortraitBooth: A Versatile Portrait Model for Fast Identity-Preserved PersonalizationXu Peng, Junwei Zhu, Boyuan Jiang, Ying Tai 等CVPR 2024 · 被引用 28 次
- Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image GenerationShuang Li, Chao Deng, Hang Chen, Liqun Liu 等CVPR 2026
- Dis²Booth: Learning Image Distribution with Disentangled Features for Text-to-Image Diffusion ModelsGuanqi Ding, Chengyu Yang, Shuhui Wang, Xincheng Li 等AAAI 2025
