Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
Yeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin, Wonjong Rhee, Nojun Kwak
Abstract
In a surge of text-to-image (T2I) models and their customization methods that generate new images of a user-provided subject, current works focus on alleviating the costs incurred by a lengthy per-subject optimization. These zero-shot customization methods encode the image of a specified subject into a visual embedding which is then utilized alongside the textual embedding for diffusion guidance. The visual embedding incorporates intrinsic information about the subject, while the textual embedding provides a new context. However, the existing methods often 1) generate images with the same pose as an input image, and 2) exhibit deterioration in the subject's identity when facing a pose variation prompt. We first pin down the problem and show that redundant pose information in the visual embedding interferes with the pose indication in the textual embedding. Conversely, the textual embedding also harms the subject's identity which is tightly entangled with the pose in the visual embedding. As a remedy, we propose text-orthogonal visual embedding which effectively harmonizes with the given textual embedding. We also adopt the visual-only embedding and inject the subject's clear features utilizing a self-attention swap. Our method is both effective and robust, offering highly flexible zero-shot generation while effectively maintaining the subject's identity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb86f458-74ed-44e0-a1de-1f0655cc040fCited by top-tier papers4
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang et al.ICLR 2026 · 34 citations
- ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion PersonalizationShizhan Liu, Hao Zheng, Hang Yu, Jianguo LiICLR 2026 · 1 citation
- Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image CustomizationLiyuan Ma, Xueji Fang, Guo-Jun QiACM MM 2024
- Semantic Alignment for Pose-Invariant Identity Preserving DiffusionJiwon Kim, Seonhwa Kim, Soobin Park, Eunju Cha et al.CVPR 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- UniversalBooth: Model-Agnostic Personalized Text-To-Image GenerationSonghua Liu, Ruonan Yu, Xinchao WangICCV 2025 · 2 citations
- Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion modelsKyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo ShinNeurIPS 2024 · 13 citations
- DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image GenerationHong Chen, Yipeng Zhang, Simin Wu, Xin Wang et al.ICLR 2024 · 81 citations
- Zero-Shot Contrastive Loss for Text-Guided Diffusion Image Style TransferSerin Yang, Hyunmin Hwang, Jong Chul YeICCV 2023 · 94 citations
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
