Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement
Nuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang, Xinbo Gao
摘要
Adversarial fine-tuning methods enhance adversarial robustness via fine-tuning the pre-trained model in an adversarial training manner. However, we identify that some specific latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap between features in the last hidden layer of natural and adversarial samples. To address this issue, we propose a disentanglement-based approach to explicitly model and further remove the specific latent features. We introduce a feature disentangler to separate out the specific latent features from the features of the adversarial samples, thereby boosting robustness by eliminating the specific latent features. Besides, we align clean features in the pre-trained model with features of adversarial samples in the fine-tuned model, to benefit from the intrinsic features of natural samples. Empirical evaluations on three benchmark datasets demonstrate that our approach surpasses existing adversarial fine-tuning methods and adversarial training baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper28
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao 等ICML 2022 · 被引用 663 次
相关 Paper
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 被引用 38 次
- AutoLoRa: An Automated Robust Fine-Tuning FrameworkXilie Xu, Jingfeng Zhang, Mohan S. KankanhalliICLR 2024 · 被引用 5 次
- Towards Defending against Adversarial Examples via Attack-Invariant FeaturesDawei Zhou, Tongliang Liu, Bo Han, Nannan Wang 等ICML 2021 · 被引用 55 次
- Feature Separation and Recalibration for Adversarial RobustnessWoo Jae Kim, Yoonki Cho, Junsik Jung, Sung-Eui YoonCVPR 2023
- Distilling Robust and Non-Robust Features in Adversarial Examples by Information BottleneckJunho Kim, Byung-Kwan Lee, Yong Man RoNeurIPS 2021 · 被引用 57 次
