Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
Lue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang, Songcan Chen
摘要
Delusive attacks aim to substantially deteriorate the test accuracy of the learning model by slightly perturbing the features of correctly labeled training examples. By formalizing this malicious attack as finding the worst-case training data within a specific -Wasserstein ball, we show that minimizing adversarial risk on the perturbed data is equivalent to optimizing an upper bound of natural risk on the original data. This implies that adversarial training can serve as a principled defense against delusive attacks. Thus, the test accuracy decreased by delusive attacks can be largely recovered by adversarial training. To further understand the internal mechanism of the defense, we disclose that adversarial training can resist the delusive perturbations by preventing the learner from overly relying on non-robust features in a natural setting. Finally, we complement our theoretical findings with a set of experiments on popular benchmark datasets, which show that the defense withstands six different practical attacks. Both theoretical and empirical results vote for adversarial training when confronted with delusive adversaries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Adversarial Neuron Pruning Purifies Backdoored Deep ModelsDongxian Wu, Yisen WangNeurIPS 2021 · 被引用 441 次
- Robust Unlearnable Examples: Protecting Data Privacy Against Adversarial LearningShaopeng Fu, Fengxiang He, Yang Liu, Li Shen 等ICLR 2022 · 被引用 64 次
- Autoregressive Perturbations for Data PoisoningPedro Sandoval Segura, Vasu Singla, Jonas Geiping, Micah Goldblum 等NeurIPS 2022 · 被引用 62 次
- Image Shortcut Squeezing: Countering Perturbative Availability Poisons with CompressionZhuoran Liu, Zhengyu Zhao, Martha A. LarsonICML 2023 · 被引用 51 次
- Not All Poisons are Created Equal: Robust Training against Data PoisoningYu Yang, Tian Yu Liu, Baharan MirzasoleimanICML 2022 · 被引用 45 次
它引用的顶会 Paper62
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
相关 Paper
- Are Adversarial Examples Created Equal? A Learnable Weighted Minimax Risk for Robustness under Non-uniform AttacksHuimin Zeng, Chen Zhu, Tom Goldstein, Furong HuangAAAI 2021 · 被引用 21 次
- DICE: Domain-attack Invariant Causal Learning for Improved Data Privacy Protection and Adversarial RobustnessQibing Ren, Yiting Chen, Yichuan Mo, Qitian Wu 等KDD 2022 · 被引用 6 次
- Fundamental Tradeoffs in Distributionally Adversarial TrainingMohammad Mehrabi, Adel Javanmard, Ryan A. Rossi, Anup B. Rao 等ICML 2021 · 被引用 19 次
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 被引用 15 次
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
