GNN Explanations that do not Explain and How to find Them
Steve Azzolin, Stefano Teso, Bruno Lepri, Andrea Passerini, Sagar Malhotra
Abstract
Explanations provided by Self-explainable Graph Neural Networks (SE-GNNs) are fundamental for understanding the model's inner workings and for identifying potential misuse of sensitive attributes. Although recent works have highlighted that these explanations can be suboptimal and potentially misleading, a characterization of their failure cases is unavailable. In this work, we identify a critical failure of SE-GNN explanations: explanations can be unambiguously unrelated to how the SE-GNNs infer labels. We show that, on the one hand, many SE-GNNs can achieve optimal true risk while producing these degenerate explanations, and on the other, most faithfulness metrics can fail to identify these failure modes. Our empirical analysis reveals that degenerate explanations can be maliciously planted (allowing an attacker to hide the use of sensitive attributes) and can also emerge naturally, highlighting the need for reliable auditing. To address this, we introduce a novel faithfulness metric that reliably marks degenerate explanations as unfaithful, in both malicious and natural settings. Our code is available on GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39d75f9b-2ac4-4cfe-8a65-134e2728f00dBuilds on44
- Parameterized Explainer for Graph Neural NetworkDongsheng Luo, Wei Cheng, Dongkuan Xu, Wenchao Yu et al.NeurIPS 2020 · 888 citations
- On Explainability of Graph Neural Networks via Subgraph ExplorationsHao Yuan, Haiyang Yu, Jie Wang, Kang Li et al.ICML 2021 · 498 citations
- Interpretable and Generalizable Graph Learning via Stochastic Attention MechanismSiqi Miao, Mia Liu, Pan LiICML 2022 · 288 citations
- Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsYongqiang Chen, Yonggang Zhang, Yatao Bian, Han Yang et al.NeurIPS 2022 · 246 citations
- Graph Information Bottleneck for Subgraph RecognitionJunchi Yu, Tingyang Xu, Yu Rong, Yatao Bian et al.ICLR 2021 · 200 citations
Related papers
- Reconsidering Faithfulness in Regular, Self-Explainable and Domain Invariant GNNsSteve Azzolin, Antonio Longa, Stefano Teso, Andrea PasseriniICLR 2025
- Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNsWenxin Tai, Ting Zhong, Goce Trajcevski, Fan ZhouICLR 2026
- Redundancy Undermines the Trustworthiness of Self-Interpretable GNNsWenxin Tai, Ting Zhong, Goce Trajcevski, Fan ZhouICML 2025
- Jointly Attacking Graph Neural Network and its ExplanationsWenqi Fan, Han Xu, Wei Jin, Xiaorui Liu et al.ICDE 2023 · 23 citations
- DEGREE: Decomposition Based Explanation for Graph Neural NetworksQizhang Feng, Ninghao Liu, Fan Yang, Ruixiang Tang et al.ICLR 2022 · 33 citations
