X2i: Seamless Integration of Multimodal Understanding Into Diffusion Transformer Via Attention Distillation
Jian Ma, Qirong Peng, Xu Guo, Chen Chen, Haonan Lu, Zhenyu Yang
Abstract
Text-to-image (T2I) models are well known for their ability to produce highly realistic images, while multimodal large language models (MLLMs) are renowned for their proficiency in understanding and integrating multiple modalities. However, currently there is no straightforward and efficient framework to transfer the multimodal comprehension abilities of MLLMs to T2I models to enable them to understand multimodal inputs. In this paper, we propose the X2I framework, which endows Diffusion Transformer (DiT) models with the capability to comprehend various modalities, including multilingual text, screenshot documents, images, videos, and audio. X2I is trained on a 100 K English corpus in 160 GPU hours. Building on the DiT teacher model, we adopt an innovative distillation method to extract the inference capabilities of the teacher model and design a lightweight AlignNet structure to serve as an intermediate bridge. Compared to the teacher model, X2I shows a decrease in performance degradation of less than 1 % while gaining various multimodal understanding abilities, including multilingual to image, image to image, image-text to image, video to image, audio to image, and utilizing creative fusion to enhance imagery. Furthermore, it is applicable for LoRA training in the context of image-text to image generation, filling a void in the industry in this area. We further design a simple LightControl to enhance the fidelity of instructional image editing. Finally, extensive experiments demonstrate the effectiveness, efficiency, multifunctionality, and transferability of our X2I. The open-source code and checkpoints for X2I can be found at the following link: https://github.com/OPPO-Mente-Lab/X2I.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Does FLUX Already Know How to Perform Physically Plausible Image Composition?Shilin Lu, Zhuming Lian, Zihan Zhou, Shaocong Zhang et al.ICLR 2026 · 34 citations
- X2Edit: Revisiting Arbitrary-Instruction Image Editing Through Self-Constructed Data and Task-Aware Representation LearningJian Ma, Xujie Zhu, Zihao Pan, Qirong Peng et al.AAAI 2026 · 15 citations
- Neodragon: Mobile Video Generation Using Diffusion TransformerAnimesh Karnewar, Denis Korzhenkov, Ioannis Lelekas, Noor Fathima et al.ICLR 2026 · 10 citations
- From Scale to Speed: Adaptive Test-Time Scaling for Image EditingXiangyan Qu, Zhenlong Yuan, Jing Tang, Rui Chen et al.CVPR 2026 · 8 citations
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion TransformersJian Ma, Qirong Peng, Xujie Zhu, Peixing Xie et al.CVPR 2026 · 7 citations
Builds on41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Dual Diffusion for Unified Image Generation and UnderstandingZijie Li, Henry Li, Yichun Shi, Amir Barati Farimani et al.CVPR 2025
- AltDiffusion: A Multilingual Text-to-Image Diffusion ModelFulong Ye, Guang Liu, Xinya Wu, Ledell WuAAAI 2024 · 54 citations
- One Transformer Fits All Distributions in Multi-Modal Diffusion at ScaleFan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li et al.ICML 2023 · 236 citations
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and PerceptionXinyang Song, Libin Wang, Weining Wang, Shaozhen Liu et al.AAAI 2026
- MultiFusion: Fusing Pre-Trained Models for Multi-Lingual, Multi-Modal Image GenerationMarco Bellagente, Manuel Brack, Hannah Teufel, Felix Friedrich et al.NeurIPS 2023 · 31 citations
