Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Bingda Tang, Boyang Zheng, Sayak Paul, Saining Xie
Abstract
This paper does not describe a new method; instead, it provides a thorough exploration of an important yet understudied design space related to recent advances in text-to-image synthesis-specifically, the deep fusion of large language models (LLMs) with diffusion transformers (DiTs) for multimodal generation. Previous studies mainly focused on overall system performance rather than detailed comparisons with alternative methods, and key design details and training recipes were often left undisclosed. These gaps create uncertainty about the real potential of this approach. To fill these gaps, we conduct an empirical study on text-to-image generation, performing controlled comparisons with established baselines, analyzing important design choices, and providing a clear, reproducible recipe for training at scale. We hope this work offers meaningful data points and practical guidelines for future research in multimodal generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cca5282-b7da-45d4-991c-7542c7dfc88fCited by top-tier papers4
- Uniform Discrete Diffusion with Metric Path for Video GenerationHaoge Deng, Ting Pan, Fan Zhang, Yang Liu et al.ICLR 2026 · 12 citations
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation ModelingYuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu et al.CVPR 2026 · 9 citations
- Scaling Sequence-to-Sequence Generative Neural RenderingShikun Liu, Kam Woh Ng, Wonbong Jang, Jiadong Guo et al.ICLR 2026 · 9 citations
- Mixture of States: Routing Token-Level Dynamics for Multimodal GenerationHaozhe Liu, Ding Liu, Mingchen Zhuge, Zijian Zhou et al.CVPR 2026 · 2 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu et al.ICML 2025
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion ModelsBingqi Ma, Zhuofan Zong, Guanglu Song, Hongsheng Li et al.NeurIPS 2024 · 57 citations
- X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention DistillationJian Ma, Qirong Peng, Xu Guo, Chen Chen et al.ICCV 2025 · 1 citation
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan et al.ICLR 2024 · 61 citations
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsSicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer et al.ICCV 2025
