Splitting the Difference on Adversarial Training
Matan Levi, Aryeh Kontorovich
摘要
The existence of adversarial examples points to a basic weakness of deep neural networks. One of the most effective defenses against such examples, adversarial training, entails training models with some degree of robustness, usually at the expense of a degraded natural accuracy. Most adversarial training methods aim to learn a model that finds, for each class, a common decision boundary encompassing both the clean and perturbed examples. In this work, we take a fundamentally different approach by treating the perturbed examples of each class as a separate class to be learned, effectively splitting each class into two classes:"clean"and"adversarial."This split doubles the number of classes to be learned, but at the same time considerably simplifies the decision boundaries. We provide a theoretical plausibility argument that sheds some light on the conditions under which our approach can be expected to be beneficial. Likewise, we empirically demonstrate that our method learns robust models while attaining optimal or near-optimal natural accuracy, e.g., on CIFAR-10 we obtain near-optimal natural accuracy of alongside significant robustness across multiple tasks. The ability to achieve such near-optimal natural accuracy, while maintaining a significant level of robustness, makes our method applicable to real-world applications where natural accuracy is at a premium. As a whole, our main contribution is a general method that confers a significant level of robustness upon classifiers with only minor or negligible degradation of their natural accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CyberPal.AI: Empowering LLMs with Expert-Driven Cybersecurity InstructionsMatan Levi, Yair Allouche, Daniel Ohayon, Anton PuzanovAAAI 2025 · 被引用 17 次
- MIMIR: Masked Image Modeling for Mutual Information-based Adversarial RobustnessXiaoyun Xu, Shujian Yu, Zhuoran Liu, Stjepan PicekNDSS 2026 · 被引用 12 次
- Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language ModelsJun Feng, Shuhong Wu, Hong Sun, Pengfei Zhang 等AAAI 2026
- Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property RefinementJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangCCS 2025
它引用的顶会 Paper31
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
相关 Paper
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Reducing Excessive Margin to Achieve a Better Accuracy vs. Robustness Trade-offRahul Rade, Seyed-Mohsen Moosavi-DezfooliICLR 2022 · 被引用 166 次
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu 等ICCV 2023 · 被引用 6 次
- Generalist: Decoupling Natural and Robust GeneralizationHongjun Wang, Yisen WangCVPR 2023
- The Enemy of My Enemy is My Friend: Exploring Inverse Adversaries for Improving Adversarial TrainingJunhao Dong, Seyed-Mohsen Moosavi-Dezfooli, Jianhuang Lai, Xiaohua XieCVPR 2023
