Securely Fine-tuning Pre-trained Encoders Against Adversarial Examples
Ziqi Zhou, Minghui Li, Wei Liu, Shengshan Hu, Yechao Zhang, Wei Wan, Lulu Xue, Leo Yu Zhang, Dezhong Yao, Hai Jin
摘要
With the evolution of self-supervised learning, the pre-training paradigm has emerged as a predominant solution within the deep learning landscape. Model providers furnish pre-trained encoders designed to function as versatile feature extractors, enabling downstream users to harness the benefits of expansive models with minimal effort through fine-tuning. Nevertheless, recent works have exposed a vulnerability in pre-trained encoders, highlighting their susceptibility to downstream-agnostic adversarial examples (DAEs) meticulously crafted by attackers. The lingering question pertains to the feasibility of fortifying the robustness of downstream models against DAEs, particularly in scenarios where the pre-trained encoders are publicly accessible to the attackers.In this paper, we initially delve into existing defensive mechanisms against adversarial examples within the pre-training paradigm. Our findings reveal that the failure of current defenses stems from the domain shift between pre-training data and downstream tasks, as well as the sensitivity of encoder parameters. In response to these challenges, we propose Genetic Evolution-Nurtured Adversarial Fine-tuning (Gen-AF), a two-stage adversarial fine-tuning approach aimed at enhancing the robustness of downstream models. Gen-AF employs a genetic-directed dual-track adversarial fine-tuning strategy in its first stage to effectively inherit the pre-trained encoder. This involves optimizing the pre-trained encoder and classifier separately while incorporating genetic regularization to preserve the model’s topology. In the second stage, Gen-AF assesses the robust sensitivity of each layer and creates a dictionary, based on which the top-k robust redundant layers are selected with the remaining layers held fixed. Upon this foundation, we conduct evolutionary adaptability fine-tuning to further enhance the model’s generalizability. Our extensive experiments, conducted across ten self-supervised training methods and six datasets, demonstrate that Gen-AF attains high testing accuracy and robust testing accuracy against state-of-the-art DAEs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DarkSAM: Fooling Segment Anything Model to Segment NothingZiqi Zhou, Yufei Song, Minghui Li, Shengshan Hu 等NeurIPS 2024 · 被引用 44 次
- Transferable Adversarial Attacks on SAM and Its Downstream ModelsSong Xia, Wenhan Yang, Yi Yu, Xun Lin 等NeurIPS 2024 · 被引用 29 次
- AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied AgentsYichen Wang, Hangtao Zhang, Hewen Pan, Ziqi Zhou 等NeurIPS 2025 · 被引用 27 次
- Vanish into Thin Air: Cross-prompt Universal Adversarial Attacks for SAM2Ziqi Zhou, Yifan Hu, Yufei Song, Zijing Li 等NeurIPS 2025 · 被引用 17 次
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
相关 Paper
- Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained EncodersZhuxin Lei, Ziyuan Yang, Yi ZhangICLR 2026
- Downstream-agnostic Adversarial ExamplesZiqi Zhou, Shengshan Hu, Ruizhi Zhao, Qian Wang 等ICCV 2023 · 被引用 45 次
- REaaS: Enabling Adversarially Robust Downstream Classifiers via Robust Encoder as a ServiceWenjie Qu, Jinyuan Jia, Neil Zhenqiang GongNDSS 2023
- Pre-trained Adversarial PerturbationsYuanhao Ban, Yinpeng DongNeurIPS 2022 · 被引用 37 次
- TWINS: A Fine-Tuning Framework for Improved Transferability of Adversarial Robustness and GeneralizationZiquan Liu, Yi Xu, Xiangyang Ji, Antoni B. ChanCVPR 2023
