Towards Modeling Uncertainties of Self-Explaining Neural Networks via Conformal Prediction
Wei Qian, Chenxu Zhao, Yangyi Li, Fenglong Ma, Chao Zhang, Mengdi Huai
Abstract
Despite the recent progress in deep neural networks (DNNs), it remains challenging to explain the predictions made by DNNs. Existing explanation methods for DNNs mainly focus on post-hoc explanations where another explanatory model is employed to provide explanations. The fact that post-hoc methods can fail to reveal the actual original reasoning process of DNNs raises the need to build DNNs with built-in interpretability. Motivated by this, many self-explaining neural networks have been proposed to generate not only accurate predictions but also clear and intuitive insights into why a particular decision was made. However, existing self-explaining networks are limited in providing distribution-free uncertainty quantification for the two simultaneously generated prediction outcomes (i.e., a sample's final prediction and its corresponding explanations for interpreting that prediction). Importantly, they also fail to establish a connection between the confidence values assigned to the generated explanations in the interpretation layer and those allocated to the final predictions in the ultimate prediction layer. To tackle the aforementioned challenges, in this paper, we design a novel uncertainty modeling framework for self-explaining networks, which not only demonstrates strong distribution-free uncertainty modeling performance for the generated explanations in the interpretation layer but also excels in producing efficient and effective prediction sets for the final predictions based on the informative high-level basis explanations. We perform the theoretical analysis for the proposed framework. Extensive experimental evaluation demonstrates the effectiveness of the proposed uncertainty framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f4b4c468-8f7e-4803-8705-090f2e90d4e1Cited by top-tier papers5
- Data Poisoning Attacks against Conformal PredictionYangyi Li, Aobo Chen, Wei Qian, Chenxu Zhao et al.ICML 2024 · 10 citations
- Membership Inference Attacks With False Discovery Rate ControlChenxu Zhao, Wei Qian, Aobo Chen, Mengdi HuaiICCV 2025 · 2 citations
- Quantifying and Understanding Uncertainty in Large Reasoning ModelsYangyi Li, Chenxu Zhao, Mengdi HuaiACL 2026
- Model Uncertainty Quantification by Conformal Prediction in Continual LearningRui Gao, Weiwei LiuICML 2025
- AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing AttacksVan Nguyen, Tingmin Wu, Xingliang Yuan, Marthie Grobler et al.ICLR 2025
Builds on29
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Adaptive Conformal Inference Under Distribution ShiftIsaac Gibbs, Emmanuel J. CandèsNeurIPS 2021 · 665 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
Related papers
- Enhancing Interpretability for Vision Models via Shapley Value OptimizationKanglong Fan, Yunqiao Yang, Chen MaAAAI 2026
- A Framework for Learning Ante-hoc Explainable Models via ConceptsAnirban Sarkar, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. BalasubramanianCVPR 2022 · 40 citations
- Estimating Regression Predictive Distributions with Sample NetworksAli Harakeh, Jordan Sir Kwang Hu, Naiqing Guan, Steven L. Waslander et al.AAAI 2023 · 5 citations
- Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution GuidanceJung-Ho Hong, Woo-Jeoung Nam, Kyu-Sung Jeon, Seong-Whan LeeAAAI 2023 · 3 citations
- Self-explaining deep models with logic rule reasoningSeungeon Lee, Xiting Wang, Sungwon Han, Xiaoyuan Yi et al.NeurIPS 2022 · 27 citations
