Lipschitz Continuity Guided Knowledge Distillation
Yuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie, Yan Yan
Abstract
Knowledge distillation has become one of the most important model compression techniques by distilling knowledge from larger teacher networks to smaller student ones. Although great success has been achieved by prior distillation methods via delicately designing various types of knowledge, they overlook the functional properties of neural networks, which makes the process of applying those techniques to new tasks unreliable and non-trivial. To alleviate such problem, in this paper, we initially leverage Lipschitz continuity to better represent the functional characteristic of neural networks and guide the knowledge distillation process. In particular, we propose a novel Lipschitz Continuity Guided Knowledge Distillation framework to faithfully distill knowledge by minimizing the distance between two neural networks' Lipschitz constants, which enables teacher networks to better regularize student networks and improve the corresponding performance. We derive an explainable approximation algorithm with an explicit theoretical derivation to address the NP-hard problem of calculating the Lipschitz constant. Experimental results have shown that our method outperforms other benchmarks over several knowledge distillation tasks (e.g., classification, segmentation and object detection) on CIFAR-100, ImageNet, and PASCAL VOC datasets. Our code is available at https://github.com/42Shawn/LONDON/ tree/master .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c33161e-d918-4c0d-b226-a6ff547e27e0Cited by top-tier papers8
- On the Effectiveness of Lipschitz-Driven Rehearsal in Continual LearningLorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato et al.NeurIPS 2022 · 64 citations
- HeteFedRec: Federated Recommender Systems with Model HeterogeneityWei Yuan, Liang Qu, Lizhen Cui, Yongxin Tong et al.ICDE 2024 · 35 citations
- Enhancing Post-Training Quantization Calibration Through Contrastive LearningYuzhang Shang, Gaowen Liu, Ramana Rao Kompella, Yan YanCVPR 2024 · 10 citations
- Causal-DFQ: Causality Guided Data-free Network QuantizationYuzhang Shang, Bingxin Xu, Gaowen Liu, Ramana Rao Kompella et al.ICCV 2023 · 8 citations
- Efficient Multitask Dense Predictor via BinarizationYuzhang Shang, Dan Xu, Gaowen Liu, Ramana Rao Kompella et al.CVPR 2024 · 6 citations
Builds on4
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li et al.ICLR 2020 · 83 citations
- AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural NetworksJiancheng Lyu, Shuai Zhang, Yingyong Qi, Jack XinKDD 2020 · 18 citations
Related papers
- Toward Student-oriented Teacher Network Training for Knowledge DistillationChengyu Dong, Liyuan Liu, Jingbo ShangICLR 2024 · 10 citations
- Flow-Based Knowledge Transfer for Efficient Large Model DistillationXinye Yang, Junhao Wang, Rui Li, Haosen Sun et al.AAAI 2026
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- Achieving Domain-Independent Certified Robustness via Knowledge ContinuityAlan Sun, Chiyu Ma, Kenneth Ge, Soroush VosoughiNeurIPS 2024 · 3 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
