Selectively Extracting and Injecting Visual Attributes into Text-to-Image Models
Seunghwan Choi, Jooyeol Yun, Youngdo Lee, Jaegul Choo
Abstract
Text-to-image models are increasingly utilized in design workflows, but articulating nuanced design intentions through text remains a challenge. This work proposes a method that extracts a visual attribute from a reference image and injects it directly into the generation pipeline. The method optimizes a text token to exclusively represent the target attribute using a custom training prompt and two novel embeddings: distilled embedding and residual embedding. Through this approach, a wide range of attributes can be extracted, including the shape, material, or color of an object, as well as the camera angle of the image. The method is validated on various target attributes and text prompts drawn from a newly constructed dataset. The results show that it outperforms existing approaches in selectively extracting and applying target attributes across diverse contexts. Ultimately, the proposed method enables intuitive and controllable text-to-image generation, streamlining the design process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- ITI-Gen: Inclusive Text-to-Image GenerationCheng Zhang, Xuanbai Chen, Siqi Chai, Chen Henry Wu et al.ICCV 2023 · 89 citations
- Alterbute: Editing Intrinsic Attributes of Objects in ImagesTal Reiss, Daniel Winter, Matan Cohen, Alex Rav-Acha et al.ICML 2026
- Free-Lunch Color-Texture Disentanglement for Stylized Image GenerationJiang Qin, Alexandra Gomez-Villa, Senmao Li, Shiqi Yang et al.NeurIPS 2025 · 12 citations
- ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token OptimizationMinseo Kim, Minchan Kwon, Dongyeun Lee, Yunho Jeon et al.CVPR 2026 · 1 citation
