Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral Constraints
Guanjie Chen, Xinyu Zhao, Yucheng Zhou, Xiaoye Qu, Tianlong Chen, Yu Cheng
Abstract
Diffusion Transformers (DiT) have emerged as a powerful architecture for image and video generation, offering superior quality and scalability. However, their practical application suffers from inherent dynamic feature instability, leading to error amplification during cached inference. Through systematic analysis, we identify the absence of long-range feature preservation mechanisms as the root cause of unstable feature propagation and perturbation sensitivity. To this end, we propose Skip-DiT, an image and video generative DiT variant enhanced with Long-Skip-Connections (LSCs) - the key efficiency component in U-Nets. Theoretical spectral norm and visualization analysis demonstrate how LSCs stabilize feature dynamics. Skip-DiTarchitecture and its stabilized dynamic feature enable an efficient statical caching mechanism that reuses deep features across timesteps while updating shallow components. Extensive experiments across the image and video generation tasks demonstrate that Skip-DiTachieves: (1) training acceleration and faster convergence, (2) inference acceleration with negligible quality loss and high fidelity to the original output, outperforming existing DiT caching methods across various quantitative metrics. Our findings establish Long-Skip-Connections as critical architectural components for stable and efficient diffusion transformers. Codes are provided in the .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80fefa06-1545-4332-8447-5545466b74bdCited by top-tier papers6
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement LearningGuanjie Chen, Shirui Huang, Yifu Sun, Kai Liu et al.CVPR 2026 · 17 citations
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise CachingHanshuai Cui, Zhiqing Tang, Zhifei Xu, Zhi Yao et al.ICLR 2026 · 11 citations
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait AnimationMengchao Wang, Qiang Wang, Fan Jiang, Mu XuAAAI 2026 · 4 citations
- Multimodal Large Language Models for Multi-Subject In-Context Image GenerationYucheng Zhou, Dubing Chen, Huan Zheng, Jianbing ShenACL 2026 · 2 citations
- LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language ModelsChenglin Wang, Yucheng Zhou, Shuang Chen, Tao Wang et al.ACL 2026 · 1 citation
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion ModelsZhirong Shen, Rui Huang, Jiacheng Liu, Chang Zou et al.CVPR 2026 · 1 citation
- Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion TransformersShikang Zheng, Liang Feng, Xinyu Wang, Qinming Zhou et al.AAAI 2026 · 10 citations
- Forecast the Principal, Stabilize the Residual: Subspace-Aware Feature Caching for Diffusion TransformersGuantao Chen, Shikang Zheng, Yuqi Lin, Linfeng ZhangCVPR 2026
- ScaleLong: Towards More Stable Training of Diffusion Model via Scaling Network Long Skip ConnectionZhongzhan Huang, Pan Zhou, Shuicheng Yan, Liang LinNeurIPS 2023 · 41 citations
- Revisiting Redundancy in Diffusion Transformers: A Temporal-Spatial Joint Caching Strategy for Efficient SamplingChenxi Du, Yongheng Deng, Ju Ren, Yaoxue ZhangKDD 2026
