Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis
Bingda Tang, Boyang Zheng, Sayak Paul, Saining Xie
摘要
This paper does not describe a new method; instead, it provides a thorough exploration of an important yet understudied design space related to recent advances in text-to-image synthesis-specifically, the deep fusion of large language models (LLMs) with diffusion transformers (DiTs) for multimodal generation. Previous studies mainly focused on overall system performance rather than detailed comparisons with alternative methods, and key design details and training recipes were often left undisclosed. These gaps create uncertainty about the real potential of this approach. To fill these gaps, we conduct an empirical study on text-to-image generation, performing controlled comparisons with established baselines, analyzing important design choices, and providing a clear, reproducible recipe for training at scale. We hope this work offers meaningful data points and practical guidelines for future research in multimodal generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Uniform Discrete Diffusion with Metric Path for Video GenerationHaoge Deng, Ting Pan, Fan Zhang, Yang Liu 等ICLR 2026 · 被引用 12 次
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation ModelingYuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu 等CVPR 2026 · 被引用 9 次
- Scaling Sequence-to-Sequence Generative Neural RenderingShikun Liu, Kam Woh Ng, Wonbong Jang, Jiadong Guo 等ICLR 2026 · 被引用 9 次
- Mixture of States: Routing Token-Level Dynamics for Multimodal GenerationHaozhe Liu, Ding Liu, Mingchen Zhuge, Zijian Zhou 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- Elucidating the design space of language models for image generationXuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu 等ICML 2025
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion ModelsBingqi Ma, Zhuofan Zong, Guanglu Song, Hongsheng Li 等NeurIPS 2024 · 被引用 57 次
- X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention DistillationJian Ma, Qirong Peng, Xu Guo, Chen Chen 等ICCV 2025 · 被引用 1 次
- LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed PromptsHanan Gani, Shariq Farooq Bhat, Muzammal Naseer, Salman Khan 等ICLR 2024 · 被引用 61 次
- X-Fusion: Introducing New Modality to Frozen Large Language ModelsSicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer 等ICCV 2025
