Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
Jingyuan Wang, Yufan Wu, Mingxuan Li, Xin Lin, Junjie Wu, Chao Li
摘要
While having achieved great success in rich real-life applications, deep neural network (DNN) models have long been criticized for their vulnerability to adversarial attacks. Tremendous research efforts have been dedicated to mitigating the threats of adversarial attacks, but the essential trait of adversarial examples is not yet clear, and most existing methods are yet vulnerable to hybrid attacks and suffer from counterattacks. In light of this, in this paper, we first reveal a gradient-based correlation between sensitivity analysisbased DNN interpreters and the generation process of adversarial examples, which indicates the Achilles's heel of adversarial attacks and sheds light on linking together the two long-standing challenges of DNN: fragility and unexplainability. We then propose an interpreter-based ensemble framework called X-Ensemble for robust adversary defense. X-Ensemble adopts a novel detectionrectification process and features in building multiple sub-detectors and a rectifier upon various types of interpretation information toward target classifiers. Moreover, X-Ensemble employs the Random Forests (RF) model to combine sub-detectors into an ensemble detector for adversarial hybrid attacks defense. The non-differentiable property of RF further makes it a precious choice against the counterattack of adversaries. Extensive experiments under various types of state-of-the-art attacks and diverse attack scenarios demonstrate the advantages of X-Ensemble to competitive baseline methods. CCS CONCEPTS • Computing methodologies → Neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Automated Assertion Generation via Information Retrieval and Its Integration with Deep learningHao Yu, Yiling Lou, Ke Sun, Dezhi Ran 等ICSE 2022 · 被引用 42 次
- Full Bayesian Significance Testing for Neural NetworksZehua Liu, Zimeng Li, Jingyuan Wang, Yue HeAAAI 2024 · 被引用 14 次
它引用的顶会 Paper2
相关 Paper
- A Unified, Resilient, and Explainable Adversarial Patch DetectorVishesh Kumar, Akshay AgarwalCVPR 2025
- A Self-supervised Approach for Adversarial RobustnessMuzammal Naseer, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan 等CVPR 2020
- Self-ensemble Adversarial Training for Improved RobustnessHongjun Wang, Yisen WangICLR 2022 · 被引用 61 次
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 被引用 38 次
- Adversarial Defence by Diversified Simultaneous Training of Deep EnsemblesBo Huang, Zhiwei Ke, Yi Wang, Wei Wang 等AAAI 2021 · 被引用 20 次
