Learning for Highly Faithful Explainability
Yuhan Guo, Lizhong Ding, Shihao Jia, Yanyu Ren, Pengqi Li, Jiarun Fu, Changsheng Li, Ye Yuan, Guoren Wang
Abstract
Learning to Explain is a forward-looking paradigm recently proposed in the field of explainable AI, which envisions training explainers capable of producing high-quality explanations for target models efficiently. Although existing studies have made attempts through self-supervised optimization or learning from prior explanation methods, the Learning to Explain paradigm still faces three critical challenges: 1) self-supervised objectives rely on assumptions about the target model or task, restricting their generalizability; 2) methods driven by prior explanations struggle to guarantee the quality of the supervisory signals; and 3) depending exclusively on either approach leads to poor convergence or limited explanation quality. To address these challenges, we propose a faithfulness-guided amortized explainer that 1) theoretically derives a self-supervised objective free from assumptions about the target model or task, 2) practically generates high-quality supervisory signals by deduplicating and filtering prior explanations, and 3) jointly optimizes both objectives via a dynamic weighting strategy, enabling the amortized explainer to produce more faithful explanations for complex, high-dimensional models. We re-formalize multiple well-validated faithfulness evaluation metrics within a unified notation system and theoretically prove that an explanation mapping can simultaneously achieve optimality across all these metrics. We aggregate prior explanation methods to generate high-quality supervised signals through deduplicating and faithfulness-based filtering. Our amortized explainer leverages dynamic weighting to guide optimization, initially emphasizing pattern consistency with the supervised signals for rapid convergence, and subsequently refining explanation quality by approximating the most faithful explanation mapping. Extensive experiments across various target models and image, text, and tabular tasks demonstrate that the proposed explainer consistently outperforms all prior explanation methods across all faithfulness metrics, highlighting its effectiveness and its potential to offer a systematic solution to the fundamental challenges of the Learning to Explain paradigm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49221207-ffc5-4d3c-89f7-7f643936dc93Cited by top-tier papers2
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei et al.ICML 2026
- Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human EvaluationsPengqi Li, Lizhong Ding, Zhehao Zhou, Chunhui Zhang et al.ACL 2026
Builds on6
- FastSHAP: Real-Time Shapley Value EstimationNeil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee et al.ICLR 2022 · 186 citations
- Framework for Evaluating Faithfulness of Local ExplanationsSanjoy Dasgupta, Nave Frost, Michal MoshkovitzICML 2022 · 87 citations
- Learning to Estimate Shapley Values with Vision TransformersIan Connick Covert, Chanwoo Kim, Su-In LeeICLR 2023 · 12 citations
- Generalization Bounds for Kolmogorov-Arnold Networks (KANs) and Enhanced KANs with Lower Lipschitz ComplexityPengqi Li, Lizhong Ding, Jiarun Fu, Chunhui Zhang et al.NeurIPS 2025 · 8 citations
- Learning to Explain: Generating Stable Explanations FastXuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf et al.ACL 2021
Related papers
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 4 citations
- Unified Time Series Explanations via Amortized Optimization and Instance-level Multi-Expert Knowledge DistillationViet-Hung Tran, Zichi Zhang, Ngoc Doan, Xuan Hoang Nguyen et al.ICML 2026
- Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNsWenxin Tai, Ting Zhong, Goce Trajcevski, Fan ZhouICLR 2026
- AudioGenX: Explainability on Text-to-Audio Generative ModelsHyunju Kang, Geonhee Han, Yoonjae Jeong, Hogun ParkAAAI 2025
- Soft Local Completeness: Rethinking Completeness in XAIZiv Weiss Haddad, Oren Barkan, Yehonatan Elisha, Noam KoenigsteinICCV 2025 · 2 citations
