Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
Zhengyao Lv, Tianlin Pan, Chenyang Si, Zhaoxi Chen, Wangmeng Zuo, Ziwei Liu, Kwan-Yee K. Wong
Abstract
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mechanism of MM-DiT, namely 1) the suppression of cross-modal attention due to token imbalance between visual and textual modalities and 2) the lack of timestep-aware attention weighting, which hinder the alignment. To address these issues, we propose Temperature-Adjusted Cross-modal Attention (TACA), a parameter-efficient method that dynamically rebalances multimodal interactions through temperature scaling and timestep-dependent adjustment. When combined with LoRA fine-tuning, TACA significantly enhances text-image alignment on the T2I-CompBench benchmark with minimal computational overhead. We tested TACA on state-of-the-art models like FLUX and SD3.5, demonstrating its ability to improve image-text alignment in terms of object appearance, attribute binding, and spatial relationships. Our findings highlight the importance of balancing cross-modal attention in improving semantic fidelity in text-to-image diffusion models. Our codes are publicly available at https://github.com/Vchitect/TACA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58708f34-aaec-43b9-bd04-9450f2b1bffaCited by top-tier papers7
- Layer-wise Instance Binding for Regional and Occlusion Control in Text-to-Image Diffusion TransformersRuidong Chen, Yancheng Bai, Xuanpu Zhang, Jianhao Zeng et al.CVPR 2026 · 9 citations
- Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image GenerationZihao Wang, Yuxiang Wei, Xinpeng Zhou, Tianyu Zhang et al.CVPR 2026 · 1 citation
- What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion TransformersChenyu Zhang, Lanjun Wang, Yueyang Cheng, Ruidong Chen et al.CCS 2026 · 1 citation
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion TransformersKanghyun Baek, Jaihyun Lew, Chaehun Shin, Jungbeom Lee et al.ICML 2026 · 1 citation
- DRM: Diffusion-based Reward Model With Step-wise GuidanceJaxon Zhang, Binxin Yang, Hubery Yin, Chen Li et al.CVPR 2026 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- CTCal: Rethinking Text-to-Image Diffusion Models via Cross-Timestep Self-CalibrationXiefan Guo, Xinzhu Ma, Haiyu Zhang, Di HuangCVPR 2026 · 1 citation
- Att-Adapter: a Robust and Precise Domain-Specific Multi-Attributes T2i Diffusion Adapter Via Conditional Variational AutoencoderWonwoong Cho, Yan-Ying Chen, Matthew Klenk, David I. Inouye et al.ICCV 2025
- Seg4Diff: Unveiling Open-Vocabulary Semantic Segmentation in Text-to-Image Diffusion TransformersChaehyun Kim, Heeseong Shin, Eunbeen Hong, Heeji Yoon et al.NeurIPS 2025 · 6 citations
- Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image SynthesisQiucheng Wu, Yujian Liu, Handong Zhao, Trung Bui et al.ICCV 2023 · 55 citations
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick et al.ICCV 2025 · 9 citations
