Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
Yitong Yang, Yinglin Wang, Tian Zhang, Jing Wang, Shuting He
Abstract
While text-driven diffusion models demonstrate remarkable performance in image editing, the critical components of their text embeddings remain underexplored. The ambiguity and entanglement of these embeddings pose challenges for precise editing. In this paper, we provide a comprehensive analysis of text embeddings in Stable Diffusion XL, offering three key insights: (1) aug embedding . aug embedding is obtained by combining the pooled output of the final text encoder with the timestep embeddings. https://github.com/huggingface/diffusers retains complete textual semantics but contributes minimally to image generation as it is only fused via the ResBlocks. More text information weakens its local semantics while preserving most global semantics. (2) BOS and padding embedding do not contain any semantic information. (3) EOS holds the semantic information of all words and stylistic information. Each word embedding is important and does not interfere with the semantic injection of other embeddings. Based on these insights, we propose PSP (Prompt-Softbox-Prompt), a training-free image editing method that leverages free-text embedding. PSP enables precise image editing by modifying text embeddings within the cross-attention layers and using Softbox to control the specific area for semantic injection. This technique enables the addition and replacement of objects without affecting other areas of the image. Additionally, PSP can achieve style transfer by simply replacing text embeddings. Extensive experiments show that PSP performs remarkably well in tasks such as object replacement, object addition, and style transfer. Our code is available at https://github.com/yangyt46/PSP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff9d1bd0-8988-4062-a3ad-6f364275b59bCited by top-tier papers2
- SplitFlux: Learning to Decouple Content and Style from a Single ImageYitong Yang, Yinglin Wang, Changshuo Wang, Yongjun Zhang et al.CVPR 2026 · 5 citations
- DiffStyle3D: Consistent 3D Gaussian Stylization via Attention OptimizationYitong Yang, Yinglin Wang, Xuexin Liu, Jing Wang et al.ICML 2026 · 2 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Rethinking Global Text Conditioning in Diffusion TransformersNikita Starodubcev, Daniil Pakhomov, Zongze Wu, Ilya Drobyshevskiy et al.ICLR 2026
- Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image EditingBingyan Liu, Chengyu Wang, Tingfeng Cao, Kui Jia et al.CVPR 2024
- Addressing Text Embedding Leakage in Diffusion-Based Image EditingSunung Mun, Jinhwan Nam, Sunghyun Cho, Jungseul OkICCV 2025 · 10 citations
- Zero-shot Image-to-Image TranslationGaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li et al.SIGGRAPH 2023 · 355 citations
- Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion ModelsSenmao Li, Joost van de Weijer, Taihang Hu, Fahad Shahbaz Khan et al.ICLR 2024 · 56 citations
