Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
Guy Kaplan, Michael Toker, Yuval Reif, Yonatan Belinkov, Roy Schwartz
Abstract
Text-to-image generation models suffer from alignment problems, where generated images fail to accurately capture the objects and relations in the text prompt. Prior work has focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion. In this work, we investigate how semantic information is distributed across token representations in text-to-image prompts, analyzing it at two levels: (1) in-item representation-whether individual tokens represent their lexical item (i.e., a word or expression conveying a single concept), and (2) cross-item interaction-whether information flows between tokens of different lexical items. We use patching techniques to uncover encoding patterns, and find that information is usually concentrated in only one or two of the item's tokens; for example, in the item San Francisco's Golden Gate Bridge'', the token Gate''sufficiently captures the entire expression while the other tokens could effectively be discarded. Lexical items also tend to remain isolated; for instance, in the prompt a green dog'', the token dog''encodes no visual information about green''. However, in some cases, items do influence each other's representation, often leading to misinterpretations-e.g., in the prompt a pool by a table'', the token pool''represents a pool table''after contextualization. Our findings highlight the critical role of token-level encoding in image generation, and demonstrate that simple interventions at the encoding stage can substantially improve alignment and generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c057686-92ba-4d18-8f6c-abfd7c21c54eCited by top-tier papers1
Ask how each one uses itBuilds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Linguistic Binding in Diffusion Models: Enhancing Attribute Correspondence through Attention Map AlignmentRoyi Rassin, Eran Hirsch, Daniel Glickman, Shauli Ravfogel et al.NeurIPS 2023 · 212 citations
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
Related papers
- Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention MapsJeeyung Kim, Erfan Esmaeili, Qiang QiuCVPR 2025
- Circuit Mechanisms for Spatial Relation Generation in Diffusion TransformersBinxu Wang, Jingxuan Fan, Xu PanCVPR 2026 · 4 citations
- Dynamic Prompt Learning: Addressing Cross-Attention Leakage for Text-Based Image EditingKai Wang, Fei Yang, Shiqi Yang, Muhammad Atif Butt et al.NeurIPS 2023 · 108 citations
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingDongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang et al.NeurIPS 2024 · 75 citations
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion ModelsBowei Chen, Sai Bi, Hao Tan, He Zhang et al.ICLR 2026 · 36 citations
