Two Coupled Rejection Metrics Can Tell Adversarial Examples Apart
Tianyu Pang, Huishuai Zhang, Di He, Yinpeng Dong, Hang Su, Wei Chen, Jun Zhu, Tie-Yan Liu
摘要
Correctly classifying adversarial examples is an essential but challenging requirement for safely deploying machine learning models. As reported in RobustBench, even the state-of-the-art adversarially trained models struggle to exceed 67% robust test accuracy on CIFAR-10, which is far from practical. A complementary way towards robustness is to introduce a rejection option, allowing the model to not return predictions on uncertain inputs, where confidence is a commonly used certainty proxy. Along with this routine, we find that confidence and a rectified confidence (R-Con) can form two coupled rejection metrics, which could provably distinguish wrongly classified inputs from correctly classified ones. This intriguing property sheds light on using coupling strategies to better detect and reject adversarial examples. We evaluate our rectified rejection (RR) module on CIFAR-10, CIFAR-10-C, and CIFAR-100 under several attacks including adaptive ones, and demonstrate that the RR module is compatible with different adversarial training frameworks on improving robustness, with little extra computation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Synergy-of-Experts: Collaborate to Improve Adversarial RobustnessSen Cui, Jingfeng Zhang, Jian Liang, Bo Han 等NeurIPS 2022 · 被引用 12 次
- Stratified Adversarial Robustness with RejectionJiefeng Chen, Jayaram Raghuram, Jihye Choi, Xi Wu 等ICML 2023 · 被引用 4 次
- Sample-specific Noise Injection for Diffusion-based Adversarial PurificationYuhao Sun, Jiacheng Zhang, Zesheng Ye, Chaowei Xiao 等ICML 2025
- One Stone, Two Birds: Enhancing Adversarial Defense Through the Lens of Distributional DiscrepancyJiacheng Zhang, Benjamin I. P. Rubinstein, Jingfeng Zhang, Feng LiuICML 2025
- Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and RejectionNils Palumbo, Yang Guo, Xi Wu, Jiefeng Chen 等ICML 2024
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 被引用 2,213 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
相关 Paper
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen AttacksDavid Stutz, Matthias Hein, Bernt SchieleICML 2020 · 被引用 158 次
- MultiRobustBench: Benchmarking Robustness Against Multiple AttacksSihui Dai, Saeed Mahloujifar, Chong Xiang, Vikash Sehwag 等ICML 2023 · 被引用 11 次
- Understanding the Limitations of Conditional Generative ModelsEthan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, Richard S. ZemelICLR 2020 · 被引用 65 次
- Learnable Boundary Guided Adversarial TrainingJiequan Cui, Shu Liu, Liwei Wang, Jiaya JiaICCV 2021 · 被引用 152 次
- Relaxing Local RobustnessKlas Leino, Matt FredriksonNeurIPS 2021 · 被引用 12 次
