Towards Automating Model Explanations with Certified Robustness Guarantees
Mengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao, Aidong Zhang
摘要
Providing model explanations has gained significant popularity recently. In contrast with the traditional feature-level model explanations, concept-based explanations can provide explanations in the form of high-level human concepts. However, existing concept-based explanation methods implicitly follow a two-step procedure that involves human intervention. Specifically, they first need the human to be involved to define (or extract) the high-level concepts, and then manually compute the importance scores of these identified concepts in a post-hoc way. This laborious process requires significant human effort and resource expenditure due to manual work, which hinders their large-scale deployability. In practice, it is challenging to automatically generate the concept-based explanations without human intervention due to the subjectivity of defining the units of concept-based interpretability. In addition, due to its data-driven nature, the interpretability itself is also potentially susceptible to malicious manipulations. Hence, our goal in this paper is to free human from this tedious process, while ensuring that the generated explanations are provably robust to adversarial perturbations. We propose a novel concept-based interpretation method, which can not only automatically provide the prototype-based concept explanations but also provide certified robustness guarantees for the generated prototype-based explanations. We also conduct extensive experiments on real-world datasets to verify the desirable properties of the proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning AttacksWei Qian, Chenxu Zhao, Wei Le, Meiyi Ma 等KDD 2023 · 被引用 38 次
- Static and Sequential Malicious Attacks in the Context of Selective ForgettingChenxu Zhao, Wei Qian, Rex Ying, Mengdi HuaiNeurIPS 2023 · 被引用 30 次
- Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal PredictionWei Qian, Chenxu Zhao, Yangyi Li, Fenglong Ma 等AAAI 2024 · 被引用 14 次
- Robust Explanation for Free or At the Cost of FaithfulnessZeren Tan, Yang TianICML 2023 · 被引用 12 次
- Factor Graph-based Interpretable Neural NetworksYicong Li, Kuanjiu Zhou, Shuo Yu, Qiang Zhang 等ICLR 2025
它引用的顶会 Paper13
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex ModelsTom Heskes, Evi Sijben, Ioan Gabriel Bucur, Tom ClaassenNeurIPS 2020 · 被引用 235 次
- How Can I Explain This to You? An Empirical Study of Deep Neural Network Explanation MethodsJeya Vikranth Jeyakumar, Joseph Noor, Yu-Hsi Cheng, Luis Garcia 等NeurIPS 2020 · 被引用 173 次
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
相关 Paper
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar 等ICLR 2021 · 被引用 68 次
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 被引用 46 次
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- Robust Explanation Constraints for Neural NetworksMatthew Wicker, Juyeon Heo, Luca Costabello, Adrian WellerICLR 2023 · 被引用 3 次
- ConEx: Human-Interpretable Saliency Maps via Concept-Aware AttributionYehonatan Elisha, Oren Barkan, Ziv Haddad, Noam KoenigsteinICML 2026
