A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
Andrew Z. Wang, Songwei Ge, Tero Karras, Ming-Yu Liu, Yogesh Balaji
Abstract
Both text-to-image generation and large language models (LLMs) have made significant advancements. However, many text-to-image models still employ the somewhat outdated T5 and CLIP as their text encoders. In this work, we investigate the effectiveness of using modern decoder-only LLMs as text encoders for text-to-image diffusion models. We build a standardized training and evaluation pipeline that allows us to isolate and evaluate the effect of different text embeddings. We train a total of 27 text-to-image models with 12 different text encoders to analyze the critical aspects of LLMs that could impact text-to-image generation, including the approaches to extract embeddings, different LLMs variants, and model sizes. Our experiments reveal that the de facto way of using last-layer embeddings as conditioning leads to inferior performance. Instead, we explore embeddings from various layers and find that using layer-normalized averaging across all layers significantly improves alignment with complex prompts. Most LLMs with this conditioning outperform the baseline T5 model, showing enhanced performance in advanced visio-linguistic reasoning skills.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 698d5988-ab4e-45de-a45f-14dda1c118c1Cited by top-tier papers3
- MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiencyNicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton et al.ICML 2026 · 2 citations
- Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image GenerationBaoteng Li, Xianghao Zang, Xinran Wang, Xiangyu Na et al.CVPR 2026
- DuoGen: Towards Autonomous Interleaved Multimodal GenerationMin Shi, Xiaohui Zeng, Jiannan Huang, Yin Cui et al.CVPR 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion ModelsBingqi Ma, Zhuofan Zong, Guanglu Song, Hongsheng Li et al.NeurIPS 2024 · 57 citations
- Decoder-Only LLMs are Better Controllers for Diffusion ModelsZiyi Dong, Yao Xiao, Pengxu Wei, Liang LinACM MM 2024 · 3 citations
- Scaling Down Text Encoders of Text-to-Image Diffusion ModelsLifu Wang, Daqing Liu, Xinchen Liu, Xiaodong HeCVPR 2025
- VSC: Visual Search Compositional Text-to-Image Diffusion ModelDo Huu Dat, Nam Hyeon-Woo, Po Yuan Mao, Tae-Hyun OhICCV 2025 · 1 citation
- Diffusion Lens: Interpreting Text Encoders in Text-to-Image PipelinesMichael Toker, Hadas Orgad, Mor Ventura, Dana Arad et al.ACL 2024
