Adversarial Infidelity Learning for Model Interpretation
Jian Liang, Bing Bai, Yuren Cao, Kun Bai, Fei Wang
摘要
Model interpretation is essential in data mining and knowledge discovery. It can help understand the intrinsic model working mechanism and check if the model has undesired characteristics. A popular way of performing model interpretation is Instance-wise Feature Selection (IFS), which provides an importance score of each feature representing the data samples to explain how the model generates the specific output. In this paper, we propose a Model-agnostic Effective Efficient Direct (MEED) IFS framework for model interpretation, mitigating concerns about sanity, combinatorial shortcuts, model identifiability, and information transmission. Also, we focus on the following setting: using selected features to directly predict the output of the given model, which serves as a primary evaluation metric for model-interpretation methods. Apart from the features, we involve the output of the given model as an additional input to learn an explainer based on more accurate information. To learn the explainer, besides fidelity, we propose an Adversarial Infidelity Learning (AIL) mechanism to boost the explanation learning by screening relatively unimportant features. Through theoretical and experimental analysis, we show that our AIL mechanism can help learn the desired conditional distribution between selected features and targets. Moreover, we extend our framework by integrating efficient interpretation methods as proper priors to provide a warm start. Comprehensive empirical evaluation results are provided by quantitative metrics and human evaluation to demonstrate the effectiveness and superiority of our proposed method. Our code is publicly available online at https://github.com/langlrsw/MEED .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Towards Multi-Grained Explainability for Graph Neural NetworksXiang Wang, Ying-Xin Wu, An Zhang, Xiangnan He 等NeurIPS 2021 · 被引用 105 次
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li 等KDD 2021 · 被引用 51 次
- DEGREE: Decomposition Based Explanation for Graph Neural NetworksQizhang Feng, Ninghao Liu, Fan Yang, Ruixiang Tang 等ICLR 2022 · 被引用 33 次
- Self-explaining deep models with logic rule reasoningSeungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi 等NeurIPS 2022 · 被引用 27 次
- Stratified GNN Explanations through Sufficient ExpansionYuwen Ji, Lei Shi, Zhimeng Liu, Ge WangAAAI 2024 · 被引用 6 次
它引用的顶会 Paper1
相关 Paper
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 被引用 22 次
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
- Distribution-Based Feature Attribution for Explaining the Predictions of Any ClassifierXinpeng Li, Kai Ming TingAAAI 2026
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy ModelsJunhao Liu, Haonan Yu, Zhenyu Yan, Xin ZhangACL 2026 · 被引用 2 次
