Anti-Distillation Backdoor Attacks: Backdoors Can Really Survive in Knowledge Distillation
Yunjie Ge, Qian Wang, Baolin Zheng, Xinlu Zhuang, Qi Li, Chao Shen, Cong Wang
Abstract
Motivated by resource-limited scenarios, knowledge distillation (KD) has received growing attention, effectively and quickly producing lightweight yet high-performance student models by transferring the dark knowledge from large teacher models. However, many pre-trained teacher models are downloaded from public platforms that lack necessary vetting, posing a possible threat to knowledge distillation tasks. Unfortunately, thus far, there has been little research to consider the backdoor attack from the teacher model into student models in KD, which may pose a severe threat to its wide use. In this paper, we, for the first time, propose a novel Anti-Distillation Backdoor Attack (ADBA), in which the backdoor embedded in the public teacher model can survive the knowledge distillation process and thus be transferred to secret distilled student models. We first introduce a shadow to imitate the distillation process and adopt an optimizable trigger to transfer information to help craft the desired teacher model. Our attack is powerful and effective, which achieves 95.92%, 94.79%, and 90.19% average success rates of attacks (SRoAs) against several different structure student models on MNIST, CIFAR-10, and GTSRB, respectively. Our ADBA also performs robustly under different user distillation environments with 91.72% and 92.37% average SRoAs on MNIST and CIFAR-10, respectively. Finally, we show that the ADBA has a low overhead in the injecting process, which converges on 50 and 70 epochs on CIFAR-10 and GTSRB, respectively, while the normal training epochs of these datasets are almost 200.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9e416f3b-cea1-4800-b8d4-1a23bf6c806eCited by top-tier papers11
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka et al.S&P 2024 · 309 citations
- Label Poisoning is All You NeedRishi D. Jha, Jonathan Hayase, Sewoong OhNeurIPS 2023 · 58 citations
- Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural NetworksJiyang Guan, Jian Liang, Ran HeNeurIPS 2022 · 57 citations
- Physical Backdoor Attacks to Lane Detection Systems in Autonomous DrivingXingshuo Han, Guowen Xu, Yuan Zhou, Xuehuan Yang et al.ACM MM 2022 · 48 citations
- SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained EncodersTianshuo Cong, Xinlei He, Yang ZhangCCS 2022 · 28 citations
Related papers
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor AttackYukun Chen, Boheng Li, Yu Yuan, Leyi Qi et al.NeurIPS 2025 · 6 citations
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu et al.ICML 2023 · 16 citations
- Backdoor Attacks Against Dataset DistillationYugeng Liu, Zheng Li, Michael Backes, Yun Shen et al.NDSS 2023
- Poisoned Distillation: Injecting Backdoors into Distilled Datasets Without Raw Data AccessZiyuan Yang, Ming Yan, Yi Zhang, Joey Tianyi ZhouAAAI 2026
- Safe Distillation BoxJingwen Ye, Yining Mao, Jie Song, Xinchao Wang et al.AAAI 2022 · 14 citations
