Lipschitz Continuity Guided Knowledge Distillation
Yuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie, Yan Yan
摘要
Knowledge distillation has become one of the most important model compression techniques by distilling knowledge from larger teacher networks to smaller student ones. Although great success has been achieved by prior distillation methods via delicately designing various types of knowledge, they overlook the functional properties of neural networks, which makes the process of applying those techniques to new tasks unreliable and non-trivial. To alleviate such problem, in this paper, we initially leverage Lipschitz continuity to better represent the functional characteristic of neural networks and guide the knowledge distillation process. In particular, we propose a novel Lipschitz Continuity Guided Knowledge Distillation framework to faithfully distill knowledge by minimizing the distance between two neural networks' Lipschitz constants, which enables teacher networks to better regularize student networks and improve the corresponding performance. We derive an explainable approximation algorithm with an explicit theoretical derivation to address the NP-hard problem of calculating the Lipschitz constant. Experimental results have shown that our method outperforms other benchmarks over several knowledge distillation tasks (e.g., classification, segmentation and object detection) on CIFAR-100, ImageNet, and PASCAL VOC datasets. Our code is available at https://github.com/42Shawn/LONDON/ tree/master .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- On the Effectiveness of Lipschitz-Driven Rehearsal in Continual LearningLorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato 等NeurIPS 2022 · 被引用 64 次
- HeteFedRec: Federated Recommender Systems with Model HeterogeneityWei Yuan, Liang Qu, Lizhen Cui, Yongxin Tong 等ICDE 2024 · 被引用 35 次
- Enhancing Post-Training Quantization Calibration Through Contrastive LearningYuzhang Shang, Gaowen Liu, Ramana Rao Kompella, Yan YanCVPR 2024 · 被引用 10 次
- Causal-DFQ: Causality Guided Data-free Network QuantizationYuzhang Shang, Bingxin Xu, Gaowen Liu, Ramana Rao Kompella 等ICCV 2023 · 被引用 8 次
- Efficient Multitask Dense Predictor via BinarizationYuzhang Shang, Dan Xu, Gaowen Liu, Ramana Rao Kompella 等CVPR 2024 · 被引用 6 次
它引用的顶会 Paper4
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li 等ICLR 2020 · 被引用 83 次
- AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural NetworksJiancheng Lyu, Shuai Zhang, Yingyong Qi, Jack XinKDD 2020 · 被引用 18 次
相关 Paper
- Toward Student-oriented Teacher Network Training for Knowledge DistillationChengyu Dong, Liyuan Liu, Jingbo ShangICLR 2024 · 被引用 10 次
- Flow-Based Knowledge Transfer for Efficient Large Model DistillationXinye Yang, Junhao Wang, Rui Li, Haosen Sun 等AAAI 2026
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang 等CVPR 2022 · 被引用 213 次
- Achieving Domain-Independent Certified Robustness via Knowledge ContinuityAlan Sun, Chiyu Ma, Kenneth Ge, Soroush VosoughiNeurIPS 2024 · 被引用 3 次
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 被引用 8 次
