Reliably fast adversarial training via latent adversarial perturbation
Geon Yeong Park, Sang Wan Lee
摘要
While multi-step adversarial training is widely popular as an effective defense method against strong adversarial attacks, its computational cost is notoriously expensive, compared to standard training. Several single-step adversarial training methods have been proposed to mitigate the above-mentioned overhead cost; however, their performance is not sufficiently reliable depending on the optimization setting. To overcome such limitations, we deviate from the existing input-space-based adversarial training regime and propose a single-step latent adversarial training method (SLAT), which leverages the gradients of latent representation as the latent adversarial perturbation. We demonstrate that the ℓ1 norm of feature gradients is implicitly regularized through the adopted latent perturbation, thereby recovering local linearity and ensuring reliable performance, compared to the existing single-step adversarial training methods. Because latent perturbation is based on the gradients of the latent representations which can be obtained for free in the process of input gradients computation, the proposed method costs roughly the same time as the fast gradient sign method. Experiment results demonstrate that the proposed method, despite its structural simplicity, outperforms state-of-the-art accelerated adversarial training methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Make Some Noise: Reliable and Efficient Single-Step Adversarial TrainingPau de Jorge Aranda, Adel Bibi, Riccardo Volpi, Amartya Sanyal 等NeurIPS 2022 · 被引用 69 次
- Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples RegularizationRunqi Lin, Chaojian Yu, Tongliang LiuNeurIPS 2023 · 被引用 25 次
- Fast Adversarial Training with Smooth ConvergenceMengnan Zhao, Lihe Zhang, Yuqiu Kong, Baocai YinICCV 2023 · 被引用 16 次
- Layer-Aware Analysis of Catastrophic Overfitting: Revealing the Pseudo-Robust Shortcut DependencyRunqi Lin, Chaojian Yu, Bo Han, Hang Su 等ICML 2024 · 被引用 9 次
- Efficient local linearity regularization to overcome catastrophic overfittingElías Abad-Rocamora, Fanghui Liu, Grigorios Chrysos, Pablo M. Olmos 等ICLR 2024 · 被引用 9 次
它引用的顶会 Paper9
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 被引用 597 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Understanding and Improving Fast Adversarial TrainingMaksym Andriushchenko, Nicolas FlammarionNeurIPS 2020 · 被引用 366 次
相关 Paper
- Efficient Robust Training via Backward SmoothingJinghui Chen, Yu Cheng, Zhe Gan, Quanquan Gu 等AAAI 2022 · 被引用 46 次
- Towards Efficient and Effective Adversarial TrainingGaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, Venkatesh Babu R.NeurIPS 2021 · 被引用 89 次
- Understanding and Increasing Efficiency of Frank-Wolfe Adversarial TrainingTheodoros Tsiligkaridis, Jay RobertsCVPR 2022 · 被引用 6 次
- Understanding Catastrophic Overfitting in Single-step Adversarial TrainingHoki Kim, Woojin Lee, Jaewook LeeAAAI 2021 · 被引用 135 次
- Towards Better Robust Generalization with Shift Consistency RegularizationShufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 等ICML 2021 · 被引用 18 次
