Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based Testing
Jinwen He, Kai Chen, Guozhu Meng, Jiangshan Zhang, Congyi Li
摘要
While enjoying the great achievements brought by deep learning (DL), people are also worried about the decision made by DL models, since the high degree of non-linearity of DL models makes the decision extremely difficult to understand. Consequently, attacks such as adversarial attacks are easy to carry out, but difficult to detect and explain, which has led to a boom in the research on local explanation methods for explaining model decisions. In this paper, we evaluate the faithfulness of explanation methods and find that traditional tests on faithfulness encounter the random dominance problem, i.e., the random selection performs the best, especially for complex data. To further solve this problem, we propose three trend-based faithfulness tests and empirically demonstrate that the new trend tests can better assess faithfulness than traditional tests on image, natural language and security tasks. We implement the assessment system and evaluate ten popular explanation methods. Benefiting from the trend tests, we successfully assess the explanation methods on complex data for the first time, bringing unprecedented discoveries and inspiring future research. Downstream tasks also greatly benefit from the tests. For example, model debugging equipped with faithful explanation methods performs much better for detecting and correcting accuracy and security problems. CCS CONCEPTS • Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fool Me If You Can: On the Robustness of Binary Code Similarity Detection Models against Semantics-Preserving TransformationsJiyong Uhm, Minseok Kim, Michalis Polychronakis, Hyungjoon KooFSE 2026 · 被引用 1 次
- Achieving Interpretable DL-based Web Attack Detection through Malicious Payload LocalizationPeiyang Li, Fukun Mei, Ye Wang, Zhuotao Liu 等NDSS 2026
它引用的顶会 Paper20
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Noise or Signal: The Role of Image Backgrounds in Object RecognitionKai Yuanqing Xiao, Logan Engstrom, Andrew Ilyas, Aleksander MadryICLR 2021 · 被引用 451 次
- LEMNA: Explaining Deep Learning based Security ApplicationsWenbo Guo, Dongliang Mu, Jun Xu, Purui Su 等CCS 2018 · 被引用 336 次
- Vulnerability detection with fine-grained interpretationsYi Li, Shaohua Wang, Tien N. NguyenFSE 2021 · 被引用 283 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
相关 Paper
- NILE : Natural Language Inference with Faithful Natural Language ExplanationsSawan Kumar, Partha P. TalukdarACL 2020 · 被引用 15 次
- Rules Refine the Riddle: Global Explanation for Deep Learning-Based Anomaly Detection in Security ApplicationsDongqi Han, Zhiliang Wang, Ruitao Feng, Minghui Jin 等CCS 2024 · 被引用 3 次
- A Comparative Study of Faithfulness Metrics for Model Interpretability MethodsChun Sik Chan, Huanqi Kong, Guanqing LiangACL 2022
- Logic Traps in Evaluating Attribution ScoresYiming Ju, Yuanzhe Zhang, Zhao Yang, Zhongtao Jiang 等ACL 2022
- Measuring the (Un)Faithfulness of Concept-Based ExplanationsShubham Kumar, Narendra AhujaCVPR 2026 · 被引用 1 次
