EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
Xinwang Chen, Ning Liu, Yichen Zhu, Feifei Feng, Jian Tang
Abstract
Transformer-based Diffusion Probabilistic Models (DPMs) have shown more potential than CNN-based DPMs, yet their extensive computational requirements hinder widespread practical applications. To reduce the computation budget of transformer-based DPMs, this work proposes the Efficient Diffusion Transformer (EDT) framework. The framework includes a lightweight-design diffusion model architecture, and a training-free Attention Modulation Matrix and its alternation arrangement in EDT inspired by human-like sketching. Additionally, we propose a token relation-enhanced masking training strategy tailored explicitly for EDT to augment its token relation learning capability. Our extensive experiments demonstrate the efficacy of EDT. The EDT framework reduces training and inference costs and surpasses existing transformer-based diffusion models in image synthesis performance, thereby achieving a significant overall enhancement. With lower FID, EDT-S, EDT-B, and EDT-XL attained speed-ups of 3.93x, 2.84x, and 1.92x respectively in the training phase, and 2.29x, 2.29x, and 2.22x respectively in inference, compared to the corresponding sizes of MDTv2. The source code is released at https://github.com/xinwangChen/EDT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SparseDiT: Token Sparsification for Efficient Diffusion TransformerShuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang et al.NeurIPS 2025 · 9 citations
- SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile DeviceYushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu et al.CVPR 2025
- Human-like Abstract Visual Reasoning via Understanding and Solving Reasoning LoopXinwang Chen, Xiuxing Li, Qing Li, Ziyue Zhuang et al.CVPR 2026
- Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield PerspectiveHyunmin Cho, Woo Kyoung Han, Kyong Hwan JinICML 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Masked Diffusion Transformer is a Strong Image SynthesizerShanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng YanICCV 2023 · 290 citations
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick et al.ICCV 2025 · 9 citations
- Tread: Token Routing for Efficient Architecture-Agnostic Diffusion TrainingFelix Krause, Timy Phan, Ming Gui, Stefan Andreas Baumann et al.ICCV 2025 · 1 citation
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li et al.ICLR 2026 · 2 citations
- Dynamic Diffusion TransformerWangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang et al.ICLR 2025
