Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
Yukun Chen, Boheng Li, Yu Yuan, Leyi Qi, Yiming Li, Tianwei Zhang, Zhan Qin, Kui Ren
摘要
Knowledge distillation (KD) is a vital technique for deploying deep neural networks (DNNs) on resource-constrained devices by transferring knowledge from large teacher models to lightweight student models. While teacher models from third-party platforms may undergo security verification (e.g., backdoor detection), we uncover a novel and critical threat: distillation-conditional backdoor attacks (DCBAs). DCBA injects dormant and undetectable backdoors into teacher models, which become activated in student models via the KD process, even with clean distillation datasets. While the direct extension of existing methods is ineffective for DCBA, we implement this attack by formulating it as a bilevel optimization problem and proposing a simple yet effective method (i.e., SCAR). Specifically, the inner optimization simulates the KD process by optimizing a surrogate student model, while the outer optimization leverages outputs from this surrogate to optimize the teacher model for implanting the conditional backdoor. Our SCAR addresses this complex optimization utilizing an implicit differentiation algorithm with a pre-optimized trigger injection function. Extensive experiments across diverse datasets, model architectures, and KD techniques validate the effectiveness of our SCAR and its resistance against existing backdoor detection, highlighting a significant yet previously overlooked vulnerability in the KD process. Our code is available at https://github.com/WhitolfChen/SCAR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series ForecastingQuang Duc Nguyen, Siyuan Liang, Yiming Li, Fushuo Huo 等ICML 2026
- Modulation-Based Backdoors: Leveraging Amplitude and Frequency Patterns to Attack Speaker RecognitionHanbo Cai, Pengcheng Zhang, Yan Xiao, De Li 等AAAI 2026
它引用的顶会 Paper38
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
相关 Paper
- Anti-Distillation Backdoor Attacks: Backdoors Can Really Survive in Knowledge DistillationYunjie Ge, Qian Wang, Baolin Zheng, Xinlu Zhuang 等ACM MM 2021 · 被引用 32 次
- Poisoned Distillation: Injecting Backdoors into Distilled Datasets Without Raw Data AccessZiyuan Yang, Ming Yan, Yi Zhang, Joey Tianyi ZhouAAAI 2026
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
- Backdoor Attacks Against Dataset DistillationYugeng Liu, Zheng Li, Michael Backes, Yun Shen 等NDSS 2023
- Revisiting Data-Free Knowledge Distillation with Poisoned TeachersJunyuan Hong, Yi Zeng, Shuyang Yu, Lingjuan Lyu 等ICML 2023 · 被引用 16 次
