Adversarial Infidelity Learning for Model Interpretation
Jian Liang, Bing Bai, Yuren Cao, Kun Bai, Fei Wang
Abstract
Model interpretation is essential in data mining and knowledge discovery. It can help understand the intrinsic model working mechanism and check if the model has undesired characteristics. A popular way of performing model interpretation is Instance-wise Feature Selection (IFS), which provides an importance score of each feature representing the data samples to explain how the model generates the specific output. In this paper, we propose a Model-agnostic Effective Efficient Direct (MEED) IFS framework for model interpretation, mitigating concerns about sanity, combinatorial shortcuts, model identifiability, and information transmission. Also, we focus on the following setting: using selected features to directly predict the output of the given model, which serves as a primary evaluation metric for model-interpretation methods. Apart from the features, we involve the output of the given model as an additional input to learn an explainer based on more accurate information. To learn the explainer, besides fidelity, we propose an Adversarial Infidelity Learning (AIL) mechanism to boost the explanation learning by screening relatively unimportant features. Through theoretical and experimental analysis, we show that our AIL mechanism can help learn the desired conditional distribution between selected features and targets. Moreover, we extend our framework by integrating efficient interpretation methods as proper priors to provide a warm start. Comprehensive empirical evaluation results are provided by quantitative metrics and human evaluation to demonstrate the effectiveness and superiority of our proposed method. Our code is publicly available online at https://github.com/langlrsw/MEED .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 422cb5d6-f32b-4f0a-b82a-659d0eb29eedCited by top-tier papers7
- Towards Multi-Grained Explainability for Graph Neural NetworksXiang Wang, Ying-Xin Wu, An Zhang, Xiangnan He et al.NeurIPS 2021 · 105 citations
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li et al.KDD 2021 · 51 citations
- DEGREE: Decomposition Based Explanation for Graph Neural NetworksQizhang Feng, Ninghao Liu, Fan Yang, Ruixiang Tang et al.ICLR 2022 · 33 citations
- Self-explaining deep models with logic rule reasoningSeungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi et al.NeurIPS 2022 · 27 citations
- Stratified GNN Explanations through Sufficient ExpansionYuwen Ji, Lei Shi, Zhimeng Liu, Ge WangAAAI 2024 · 6 citations
Builds on1
Related papers
- Towards Rigorous Interpretations: a Formalisation of Feature AttributionDarius Afchar, Vincent Guigue, Romain HennequinICML 2021 · 22 citations
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
- Distribution-Based Feature Attribution for Explaining the Predictions of Any ClassifierXinpeng Li, Kai Ming TingAAAI 2026
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy ModelsJunhao Liu, Haonan Yu, Zhenyu Yan, Xin ZhangACL 2026 · 2 citations
