Undistillable: Making A Nasty Teacher That CANNOT teach students
Haoyu Ma, Tianlong Chen, Ting-Kuei Hu, Chenyu You, Xiaohui Xie, Zhangyang Wang
Abstract
Knowledge Distillation (KD) is a widely used technique to transfer knowledge from pre-trained teacher models to (usually more lightweight) student models. However, in certain situations, this technique is more of a curse than a blessing. For instance, KD poses a potential risk of exposing intellectual properties (IPs): even if a trained machine learning model is released in "black boxes" (e.g., as executable software or APIs without open-sourcing code), it can still be replicated by KD through imitating input-output behaviors. To prevent this unwanted effect of KD, this paper introduces and investigates a concept called Nasty Teacher: a specially trained teacher network that yields nearly the same performance as a normal one, but would significantly degrade the performance of student models learned by imitating it. We propose a simple yet effective algorithm to build the nasty teacher, called self-undermining knowledge distillation. Specifically, we aim to maximize the difference between the output of the nasty teacher and a normal pretrained network. Extensive experiments on several datasets demonstrate that our method is effective on both standard KD and data-free KD, providing the desirable KD-immunity to model owners for the first time. We hope our preliminary study can draw more awareness and interest in this new practical problem of both social and legal importance. Our codes and pre-trained models can be found at https://github.com/VITA-Group/Nasty-Teacher .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d498a84-67bb-4e48-9aaa-ca916363772eCited by top-tier papers11
- Analyzing the Confidentiality of Undistillable Teachers in Knowledge DistillationSouvik Kundu, Qirui Sun, Yao Fu, Massoud Pedram et al.NeurIPS 2021 · 35 citations
- SAL-ViT: Towards Latency Efficient Private Inference on ViT using Selective Attention Search with a Learnable Softmax ApproximationYuke Zhang, Dake Chen, Souvik Kundu, Chenghao Li et al.ICCV 2023 · 30 citations
- Antidistillation SamplingYash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu et al.NeurIPS 2025 · 24 citations
- Practical and Efficient Model Extraction of Sentiment Analysis APIsWeibin Wu, Jianping Zhang, Victor Junqiu Wei, Xixian Chen et al.ICSE 2023 · 10 citations
- The Effect of Optimal Self-Distillation in Noisy Gaussian Mixture ModelKaito Takanami, Takashi Takahashi, Ayaka SakataNeurIPS 2025 · 4 citations
Builds on15
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 194 citations
Related papers
- Safe Distillation BoxJingwen Ye, Yining Mao, Jie Song, Xinchao Wang et al.AAAI 2022 · 14 citations
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu et al.ICML 2023 · 16 citations
- Anti-Distillation Backdoor Attacks: Backdoors Can Really Survive in Knowledge DistillationYunjie Ge, Qian Wang, Baolin Zheng, Xinlu Zhuang et al.ACM MM 2021 · 32 citations
- Teach Less, Learn More: On the Undistillable Classes in Knowledge DistillationYichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu et al.NeurIPS 2022 · 42 citations
- Zero-Shot Knowledge Distillation from a Decision-Based Black-Box ModelZi WangICML 2021 · 56 citations
