Learning for Highly Faithful Explainability
Yuhan Guo, Lizhong Ding, Shihao Jia, Yanyu Ren, Pengqi Li, Jiarun Fu, Changsheng Li, Ye Yuan, Guoren Wang
摘要
Learning to Explain is a forward-looking paradigm recently proposed in the field of explainable AI, which envisions training explainers capable of producing high-quality explanations for target models efficiently. Although existing studies have made attempts through self-supervised optimization or learning from prior explanation methods, the Learning to Explain paradigm still faces three critical challenges: 1) self-supervised objectives rely on assumptions about the target model or task, restricting their generalizability; 2) methods driven by prior explanations struggle to guarantee the quality of the supervisory signals; and 3) depending exclusively on either approach leads to poor convergence or limited explanation quality. To address these challenges, we propose a faithfulness-guided amortized explainer that 1) theoretically derives a self-supervised objective free from assumptions about the target model or task, 2) practically generates high-quality supervisory signals by deduplicating and filtering prior explanations, and 3) jointly optimizes both objectives via a dynamic weighting strategy, enabling the amortized explainer to produce more faithful explanations for complex, high-dimensional models. We re-formalize multiple well-validated faithfulness evaluation metrics within a unified notation system and theoretically prove that an explanation mapping can simultaneously achieve optimality across all these metrics. We aggregate prior explanation methods to generate high-quality supervised signals through deduplicating and faithfulness-based filtering. Our amortized explainer leverages dynamic weighting to guide optimization, initially emphasizing pattern consistency with the supervised signals for rapid convergence, and subsequently refining explanation quality by approximating the most faithful explanation mapping. Extensive experiments across various target models and image, text, and tabular tasks demonstrate that the proposed explainer consistently outperforms all prior explanation methods across all faithfulness metrics, highlighting its effectiveness and its potential to offer a systematic solution to the fundamental challenges of the Learning to Explain paradigm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- FlowMAP: Flow Matching for Generalizable Agent PlanningJiarun Fu, Lizhong Ding, Ye Yuan, Qiuning Wei 等ICML 2026
- Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human EvaluationsPengqi Li, Lizhong Ding, Zhehao Zhou, Chunhui Zhang 等ACL 2026
它引用的顶会 Paper6
- FastSHAP: Real-Time Shapley Value EstimationNeil Jethani, Mukund Sudarshan, Ian Connick Covert, Su-In Lee 等ICLR 2022 · 被引用 186 次
- Framework for Evaluating Faithfulness of Local ExplanationsSanjoy Dasgupta, Nave Frost, Michal MoshkovitzICML 2022 · 被引用 87 次
- Learning to Estimate Shapley Values with Vision TransformersIan Connick Covert, Chanwoo Kim, Su-In LeeICLR 2023 · 被引用 12 次
- Generalization Bounds for Kolmogorov-Arnold Networks (KANs) and Enhanced KANs with Lower Lipschitz ComplexityPengqi Li, Lizhong Ding, Jiarun Fu, Chunhui Zhang 等NeurIPS 2025 · 被引用 8 次
- Learning to Explain: Generating Stable Explanations FastXuelin Situ, Ingrid Zukerman, Cécile Paris, Sameen Maruf 等ACL 2021
相关 Paper
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 被引用 4 次
- Unified Time Series Explanations via Amortized Optimization and Instance-level Multi-Expert Knowledge DistillationViet-Hung Tran, Zichi Zhang, Ngoc Doan, Xuan Hoang Nguyen 等ICML 2026
- Self-Consistency Improves the Trustworthiness of Self-Interpretable GNNsWenxin Tai, Ting Zhong, Goce Trajcevski, Fan ZhouICLR 2026
- AudioGenX: Explainability on Text-to-Audio Generative ModelsHyunju Kang, Geonhee Han, Yoonjae Jeong, Hogun ParkAAAI 2025
- Soft Local Completeness: Rethinking Completeness in XAIZiv Weiss Haddad, Oren Barkan, Yehonatan Elisha, Noam KoenigsteinICCV 2025 · 被引用 2 次
