Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
Jingyuan Wang, Yufan Wu, Mingxuan Li, Xin Lin, Junjie Wu, Chao Li
Abstract
While having achieved great success in rich real-life applications, deep neural network (DNN) models have long been criticized for their vulnerability to adversarial attacks. Tremendous research efforts have been dedicated to mitigating the threats of adversarial attacks, but the essential trait of adversarial examples is not yet clear, and most existing methods are yet vulnerable to hybrid attacks and suffer from counterattacks. In light of this, in this paper, we first reveal a gradient-based correlation between sensitivity analysisbased DNN interpreters and the generation process of adversarial examples, which indicates the Achilles's heel of adversarial attacks and sheds light on linking together the two long-standing challenges of DNN: fragility and unexplainability. We then propose an interpreter-based ensemble framework called X-Ensemble for robust adversary defense. X-Ensemble adopts a novel detectionrectification process and features in building multiple sub-detectors and a rectifier upon various types of interpretation information toward target classifiers. Moreover, X-Ensemble employs the Random Forests (RF) model to combine sub-detectors into an ensemble detector for adversarial hybrid attacks defense. The non-differentiable property of RF further makes it a precious choice against the counterattack of adversaries. Extensive experiments under various types of state-of-the-art attacks and diverse attack scenarios demonstrate the advantages of X-Ensemble to competitive baseline methods. CCS CONCEPTS • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Automated Assertion Generation via Information Retrieval and Its Integration with Deep learningHao Yu, Yiling Lou, Ke Sun, Dezhi Ran et al.ICSE 2022 · 42 citations
- Full Bayesian Significance Testing for Neural NetworksZehua Liu, Zimeng Li, Jingyuan Wang, Yue HeAAAI 2024 · 14 citations
Builds on2
Related papers
- A Unified, Resilient, and Explainable Adversarial Patch DetectorVishesh Kumar, Akshay AgarwalCVPR 2025
- A Self-supervised Approach for Adversarial RobustnessMuzammal Naseer, Salman H. Khan, Munawar Hayat, Fahad Shahbaz Khan et al.CVPR 2020
- Self-ensemble Adversarial Training for Improved RobustnessHongjun Wang, Yisen WangICLR 2022 · 61 citations
- Adversarial Robustness through Disentangled RepresentationsShuo Yang, Tianyu Guo, Yunhe Wang, Chang XuAAAI 2021 · 38 citations
- Adversarial Defence by Diversified Simultaneous Training of Deep EnsemblesBo Huang, Zhiwei Ke, Yi Wang, Wei Wang et al.AAAI 2021 · 20 citations
