SAFARI: Versatile and Efficient Evaluations for Robustness of Interpretability
Wei Huang, Xingyu Zhao, Gaojie Jin, Xiaowei Huang
摘要
Interpretability of Deep Learning (DL) is a barrier to trustworthy AI. Despite great efforts made by the Explainable AI (XAI) community, explanations lack robustnessindistinguishable input perturbations may lead to different XAI results. Thus, it is vital to assess how robust DL interpretability is, given an XAI method. In this paper, we identify several challenges that the state-of-the-art is unable to cope with collectively: i) existing metrics are not comprehensive; ii) XAI techniques are highly heterogeneous; iii) misinterpretations are normally rare events. To tackle these challenges, we introduce two black-box evaluation methods, concerning the worst-case interpretation discrepancy and a probabilistic notion of how robust in general, respectively. Genetic Algorithm (GA) with bespoke fitness function is used to solve constrained optimisation for efficient worstcase evaluation. Subset Simulation (SS), dedicated to estimate rare event probabilities, is used for evaluating overall robustness. Experiments show that the accuracy, sensitivity, and efficiency of our methods outperform the state-ofthe-arts. Finally, we demonstrate two applications of our methods: ranking robust XAI methods and selecting training schemes to improve both classification and interpretation robustness. this paper. However, as suggested in [30], we use the terms explanation/interpretation specifically for individual predictions. 2 Without loss of generality, in this paper we assume the DL model is a classifier if with no further clarification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 被引用 27 次
- TrajPAC: Towards Robustness Verification of Pedestrian Trajectory Prediction ModelsLiang Zhang, Nathaniel Xu, Pengfei Yang, Gaojie Jin 等ICCV 2023 · 被引用 13 次
- Adversarial Training for Probabilistic RobustnessYi Zhang, Yuhang Chen, Zhen Chen, Wenjie Ruan 等ICCV 2025 · 被引用 3 次
- Non-Parametric Probabilistic Robustness: A Conservative Risk Estimator under Unknown Perturbation DistributionsZheng Wang, Yi Zhang, Siddartha Khastgir, carsten maple 等ICML 2026
- Local Stability of RankingsFelix S. Campbell, Yuval MoskovitchSIGMOD 2026
它引用的顶会 Paper6
- Framework for Evaluating Faithfulness of Local ExplanationsSanjoy Dasgupta, Nave Frost, Michal MoshkovitzICML 2022 · 被引用 87 次
- Enhancing Adversarial Training with Second-Order Statistics of WeightsGaojie Jin, Xinping Yi, Wei Huang, Sven Schewe 等CVPR 2022 · 被引用 46 次
- Unfooling Perturbation-Based Post Hoc ExplainersZachariah Carmichael, Walter J. ScheirerAAAI 2023 · 被引用 18 次
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 被引用 3 次
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji 等USENIX Security 2020
相关 Paper
- One step further: evaluating interpreters using metamorphic testingMing Fan, Jiali Wei, Wuxia Jin, Zhou Xu 等ISSTA 2022 · 被引用 7 次
- F-Fidelity: A Robust Framework for Faithfulness Evaluation of Explainable AIXu Zheng, Farhad Shirani, Zhuomin Chen, Chaohao Lin 等ICLR 2025
- What Do You See?: Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural BackdoorsYi-Shan Lin, Wen-Chuan Lee, Z. Berkay CelikKDD 2021 · 被引用 62 次
- Robin: A Novel Method to Produce Robust Interpreters for Deep Learning-Based Code ClassifiersZhen Li, Ruqian Zhang, Deqing Zou, Ning Wang 等ASE 2023 · 被引用 4 次
- Corrupting Neuron Explanations of Deep Visual FeaturesDivyansh Srivastava, Tuomas P. Oikarinen, Tsui-Wei WengICCV 2023 · 被引用 3 次
