Progressively Knowledge Distillation via Re-parameterizing Diffusion Reverse Process
Xufeng Yao, Fanbin Lu, Yuechen Zhang, Xinyun Zhang, Wenqian Zhao, Bei Yu
Abstract
Knowledge distillation aims at transferring knowledge from the teacher model to the student one by aligning their distributions. Feature-level distillation often uses L2 distance or its variants as the loss function, based on the assumption that outputs follow normal distributions. This poses a significant challenge when distribution gaps are substantial since this loss function ignores the variance term. To address the problem, we propose to decompose the transfer objective into small parts and optimize it progressively. This process is inspired by diffusion models from which the noise distribution is mapped to the target distribution step by step. However, directly employing diffusion models is impractical in the distillation scenario due to its heavy reverse process. To overcome this challenge, we adopt the structural re-parameterization technique to generate multiple student features to approximate the teacher features sequentially. The multiple student features are combined linearly in inference time without extra cost. We present extensive experiments performed on various transfer scenarios, such as CNN-to-CNN and Transformer-to-CNN, that validate the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cfa68c0-5f81-4fd9-9848-7cd84c94ff41Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
Related papers
- Beyond Logits: Aligning Feature Dynamics for Effective Knowledge DistillationGuoqiang Gong, Jiaxing Wang, Jin Xu, Deping Xiang et al.ACL 2025
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You et al.NeurIPS 2023 · 125 citations
- Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic TeacherYong Guo, Shulian Zhang, Haolin Pan, Jing Liu et al.ICLR 2025
- Structural Knowledge Distillation: Tractably Distilling Information for Structured PredictorXinyu Wang, Yong Jiang, Zhaohui Yan, Zixia Jia et al.ACL 2021
- Flow-Based Knowledge Transfer for Efficient Large Model DistillationXinye Yang, Junhao Wang, Rui Li, Haosen Sun et al.AAAI 2026
