Secure Transfer Learning: Training Clean Model Against Backdoor in Pre-Trained Encoder and Downstream Dataset
Yechao Zhang, Yuxuan Zhou, Tianyu Li, Minghui Li, Shengshan Hu, Wei Luo, Leo Yu Zhang
摘要
Transfer learning from pre-trained encoders has become essential in modern machine learning, enabling efficient model adaptation across diverse tasks. However, this combination of pre-training and downstream adaptation creates an expanded attack surface, exposing models to sophisticated backdoor embedding at both the encoder and dataset levels—an area often overlooked in prior research. Additionally, the limited computational resources typically available to users of pre-trained encoders constrain the effectiveness of generic backdoor defenses compared to end-to-end training from scratch. In this work, we investigate how to mitigate potential backdoor risks in resource-constrained transfer learning scenarios. Specifically, we first conduct an exhaustive analysis of existing defense strategies, revealing that many follow a reactive workflow based on assumptions that do not scale to unknown threats, novel attack types, or different training paradigms. In response, we introduce a proactive mindset focused on identifying clean elements and propose the Trusted Core (T-Core) Bootstrapping framework, which emphasizes the importance of pinpointing trustworthy data and neurons to enhance model security. Our empirical evaluations demonstrate the effectiveness and superiority of T-Core, specifically assessing 5 encoder poisoning attacks, 7 dataset poisoning attacks, and 14 baseline defenses across 5 benchmark datasets, addressing 4 scenarios of 3 potential backdoor threats.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等NeurIPS 2021 · 被引用 503 次
- Backdoor Defense via Decoupling the Training ProcessKunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin 等ICLR 2022 · 被引用 253 次
- Demon in the Variant: Statistical Analysis of DNNs for Robust Backdoor Contamination DetectionDi Tang, XiaoFeng Wang, Haixu Tang, Kehuan ZhangUSENIX Security 2021 · 被引用 242 次
- Adversarial Unlearning of Backdoors via Implicit HypergradientYi Zeng, Si Chen, Won Park, Zhuoqing Mao 等ICLR 2022 · 被引用 235 次
相关 Paper
- Protecting Model Adaptation from Trojans in the Unlabeled DataLijun Sheng, Jian Liang, Ran He, Zilei Wang 等AAAI 2025
- DIFT: Protecting Contrastive Learning Against Data Poisoning Backdoor AttacksJiang Zhu, Yulin Jin, Qingqing Ye, Zhibiao Guo 等AAAI 2026
- REFINE: Inversion-Free Backdoor Defense via Model ReprogrammingYukun Chen, Shuo Shao, Enhao Huang, Yiming Li 等ICLR 2025
- Data Poisoning Based Backdoor Attacks to Contrastive LearningJinghuai Zhang, Hongbin Liu, Jinyuan Jia, Neil Zhenqiang GongCVPR 2024 · 被引用 12 次
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 被引用 20 次
