Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
Shuang Li, Chao Deng, Hang Chen, Liqun Liu, zhenyu hu, Te Cao, Mengge Xue, Yuan Chen, Peng Shu, Huan Yu, Jie Jiang
Abstract
Subject-Driven Text-to-Image (T2I) Generation aims to preserve a subject's identity while editing its context based on a text prompt. A core challenge in this task is the ``similarity-controllability paradox'', where enhancing textual control often degrades the subject's fidelity, and vice-versa. We argue this paradox stems from the ambiguous role of text prompts, which are often tasked with describing both the subject and the desired modifications, leading to conflicting signals for the model. To resolve this, we propose DisCo, a novel framework that first Disentangles and then re-Couples visual and textual information. First, our textual-visual decoupling module isolates the sources of information: subject identity is extracted exclusively from the reference image with the entity word of the subject, while the text prompt is simplified to contain only the modification command, where the subject refers to general pronouns, eliminating descriptive ambiguity. However, this strict separation can lead to unnatural compositions between the subject and its contexts. We address this by designing a dedicated reward signal and using reinforcement learning to seamlessly recouple the visually-defined subject and the textually-generated context. Our approach effectively resolves the paradox, enabling simultaneous high-fidelity subject preservation and precise textual control. Extensive experiments demonstrate that our method achieves state-of-the-art performance, producing highly realistic and coherent images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fca30ee9-6b62-4247-91b0-ca2a98b14f8dBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image GenerationHong Chen, Yipeng Zhang, Simin Wu, Xin Wang et al.ICLR 2024 · 81 citations
- RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image CustomizationMengqi Huang, Zhendong Mao, Mingcong Liu, Qian He et al.CVPR 2024
- Resolving the Identity Crisis in Text-to-Image GenerationShubhankar Borse, Farzad Farhadzadeh, Munawar Hayat, Fatih PorikliCVPR 2026 · 2 citations
- Training-Free Consistent Text-to-Image GenerationYoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten et al.SIGGRAPH 2024 · 57 citations
- CoDi: Subject-Consistent and Pose-Diverse Text-to-Image GenerationZhanxin Gao, Beier Zhu, Liangyao, Jian Yang et al.ICLR 2026 · 1 citation
