Elastic Diffusion Transformer
Jiangshan Wang, Zeqiang Lai, Jiarui Chen, Jiayi Guo, Hang Guo, Xiu Li, Xiangyu Yue, Chunchao Guo
Abstract
Diffusion Transformers (DiT) have demonstrated remarkable generative capabilities but remain highly computationally expensive. Previous acceleration methods, such as pruning and distillation, typically rely on a fixed computational capacity, leading to insufficient acceleration and degraded generation quality. To address this limitation, we propose Elastic Diffusion Transformer (E-DiT), an adaptive acceleration framework for DiT that effectively improves efficiency while maintaining generation quality. Specifically, we observe that the generative process of DiT exhibits substantial sparsity (i.e., some computations can be skipped with minimal impact on quality), and this sparsity varies significantly across samples. Motivated by this observation, E-DiT equips each DiT block with a lightweight router that dynamically identifies sample-dependent sparsity from the input latent. Each router adaptively determines whether the corresponding block can be skipped. If the block is not skipped, the router then predicts the optimal MLP width reduction ratio within the block. During inference, we further introduce a block-level feature caching mechanism that leverages router predictions to eliminate redundant computations in a training-free manner. Extensive experiments across 2D image (Qwen-Image and FLUX) and 3D asset (Hunyuan3D-3.0) demonstrate the effectiveness of E-DiT, achieving up to ∼2× speedup with negligible loss in generation quality. Code will be available at https://github.com/ wangjiangshan0725/Elastic-DiT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Principled RL for Flow Matching Emerges from the Chunk-level Policy OptimizationYifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du et al.ICML 2026 · 12 citations
- Photon: Speedup Volume Understanding with Efficient Multimodal Large Language ModelsChengyu Fang, Heng Guo, Zheng Jiang, Chunming He et al.ICLR 2026 · 10 citations
- PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video GenerationJiangshan Wang, Kang Zhao, Jiayi Guo, Jiayu Wang et al.ICLR 2026 · 6 citations
- Embedding-perturbed Exploration Preference Optimization for Flow ModelsSujie Hu, Chubin Chen, Jiashu Zhu, Jiahong Wu et al.ICML 2026
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li et al.ICLR 2026 · 2 citations
- ScalingCache: Extreme Acceleration of DiTs through Difference Scaling and Dynamic Interval CachingLihui Gu, Jingbin He, Lianghao Su, Kang He et al.ICLR 2026
- SparseDiT: Token Sparsification for Efficient Diffusion TransformerShuning Chang, Pichao Wang, Jiasheng Tang, Fan Wang et al.NeurIPS 2025 · 9 citations
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion TransformersMoayed Haji Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park et al.CVPR 2026 · 4 citations
- Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion TransformersGuantao Chen, Shikang Zheng, Yuqi Lin, Linfeng ZhangCVPR 2026
