EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
Xinwang Chen, Ning Liu, Yichen Zhu, Feifei Feng, Jian Tang
摘要
Transformer-based Diffusion Probabilistic Models (DPMs) have shown more potential than CNN-based DPMs, yet their extensive computational requirements hinder widespread practical applications. To reduce the computation budget of transformer-based DPMs, this work proposes the Efficient Diffusion Transformer (EDT) framework. The framework includes a lightweight-design diffusion model architecture, and a training-free Attention Modulation Matrix and its alternation arrangement in EDT inspired by human-like sketching. Additionally, we propose a token relation-enhanced masking training strategy tailored explicitly for EDT to augment its token relation learning capability. Our extensive experiments demonstrate the efficacy of EDT. The EDT framework reduces training and inference costs and surpasses existing transformer-based diffusion models in image synthesis performance, thereby achieving a significant overall enhancement. With lower FID, EDT-S, EDT-B, and EDT-XL attained speed-ups of 3.93x, 2.84x, and 1.92x respectively in the training phase, and 2.29x, 2.29x, and 2.22x respectively in inference, compared to the corresponding sizes of MDTv2. The source code is released at https://github.com/xinwangChen/EDT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SparseDiT: Token Sparsification for Efficient Diffusion TransformerShuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang 等NeurIPS 2025 · 被引用 9 次
- SnapGen-V: Generating a Five-Second Video within Five Seconds on a Mobile DeviceYushu Wu, Zhixing Zhang, Yanyu Li, Yanwu Xu 等CVPR 2025
- Human-like Abstract Visual Reasoning via Understanding and Solving Reasoning LoopXinwang Chen, Xiuxing Li, Qing Li, Ziyue Zhuang 等CVPR 2026
- Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield PerspectiveHyunmin Cho, Woo Kyoung Han, Kyong Hwan JinICML 2026
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
相关 Paper
- Masked Diffusion Transformer is a Strong Image SynthesizerShanghua Gao, Pan Zhou, Ming-Ming Cheng, Shuicheng YanICCV 2023 · 被引用 290 次
- Edit: Efficient Diffusion Transformers with Linear Compressed AttentionPhilipp Becker, Abhinav Mehrotra, Ruchika Chavhan, Malcolm Chadwick 等ICCV 2025 · 被引用 9 次
- Tread: Token Routing for Efficient Architecture-Agnostic Diffusion TrainingFelix Krause, Timy Phan, Ming Gui, Stefan Andreas Baumann 等ICCV 2025 · 被引用 1 次
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li 等ICLR 2026 · 被引用 2 次
- Dynamic Diffusion TransformerWangbo Zhao, Yizeng Han, Jiasheng Tang, Kai Wang 等ICLR 2025
