TinyFusion: Diffusion Transformers Learned Shallow
Gongfan Fang, Kunjun Li, Xinyin Ma, Xinchao Wang
Abstract
Diffusion Transformers have demonstrated remarkable capabilities in image generation but often come with excessive parameterization, resulting in considerable inference overhead in real-world applications. In this work, we present TinyFusion, a depth pruning method designed to remove redundant layers from diffusion transformers via endto-end learning. The core principle of our approach is to create a pruned model with high recoverability, allowing it to regain strong performance after fine-tuning. To accomplish this, we introduce a differentiable sampling technique to make pruning learnable, paired with a co-optimized parameter to simulate future fine-tuning. While prior works focus on minimizing loss or error after pruning, our method explicitly models and optimizes the post-fine-tuning performance of pruned models. Experimental results indicate that this learnable paradigm offers substantial benefits for layer pruning of diffusion transformers, surpassing existing importance-based and error-based methods. Additionally, TinyFusion exhibits strong generalization across diverse architectures, such as DiTs, MARs, and SiTs. Experiments with DiT-XL show that TinyFusion can craft a shallow diffusion transformer at less than 7% of the pretraining cost, achieving a 2× speedup with an FID score of 2.86, outperforming competitors with comparable efficiency. Code is available at https://github.com/ VainF/TinyFusion
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang et al.ICLR 2026 · 26 citations
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsChaofan Gan, Yuanpeng Tu, Xi Chen, Tieyuan Chen et al.NeurIPS 2025 · 22 citations
- QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video GenerationJunyi Wu, Zhiteng Li, Zheng Hui, Yulun Zhang et al.ICCV 2025 · 20 citations
- Provable Separations between Memorization and Generalization in Diffusion ModelsZeqi Ye, Qijie Zhu, Molei Tao, Minshuo ChenICLR 2026 · 15 citations
- Exploring Diffusion Transformer Designs via GraftingKeshigeyan Chandrasegaran, Michael Poli, Daniel Y. Fu, Dongjun Kim et al.NeurIPS 2025 · 14 citations
Builds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Pluggable Pruning with Contiguous Layer Distillation for Diffusion TransformersJian Ma, Qirong Peng, Xujie Zhu, Peixing Xie et al.CVPR 2026 · 7 citations
- SPRINT: Sparse-Dense Residual Fusion for Efficient Diffusion TransformersDogyun Park, Moayed Haji-Ali, Yanyu Li, Willi Menapace et al.ICLR 2026 · 6 citations
- DiffSparse: Accelerating Diffusion Transformers with Learned Token SparsityHaowei Zhu, Ji Liu, Ziqiong Liu, Dong Li et al.ICLR 2026 · 2 citations
- MosaicDiff: Training-free Structural Pruning for Diffusion Model Acceleration Reflecting Pretraining DynamicsBowei Guo, Shengkun Tang, Cong Zeng, Zhiqiang ShenICCV 2025 · 2 citations
- FastFLUX: Pruning FLUX with Block-wise Replacement and Sandwich TrainingFuhan Cai, Yong Guo, Jie Li, Wenbo Li et al.AAAI 2026 · 2 citations
