A Learning Theoretic Perspective on Local Explainability
Jeffrey Li, Vaishnavh Nagarajan, Gregory Plumb, Ameet Talwalkar
Abstract
In this paper, we explore connections between interpretable machine learning and learning theory through the lens of local approximation explanations. First, we tackle the traditional problem of performance generalization and bound the test-time accuracy of a model using a notion of how locally explainable it is. Second, we explore the novel problem of explanation generalization which is an important concern for a growing class of finite sample-based local approximation explanations. Finally, we validate our theoretical results empirically and show that they reflect what can be seen in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d544464-6837-4723-9f7c-41f0af57086cCited by top-tier papers6
- On the Impact of Knowledge Distillation for Model InterpretabilityHyeongrok Han, Siwon Kim, Hyun-Soo Choi, Sungroh YoonICML 2023 · 13 citations
- Learning with Explanation ConstraintsRattana Pukdee, Dylan Sam, J. Zico Kolter, Maria-Florina Balcan et al.NeurIPS 2023 · 11 citations
- Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace AdjustmentsYifan Zhang, Tianle Ren, Fei Wang, Brian Y. LimCHI 2026 · 1 citation
- Finding Safe Zones of Markov Decision Processes PoliciesLee Cohen, Yishay Mansour, Michal MoshkovitzNeurIPS 2023 · 1 citation
- Efficient and Accurate Explanation Estimation with Distribution CompressionHubert Baniecki, Giuseppe Casalicchio, Bernd Bischl, Przemyslaw BiecekICLR 2025
Builds on1
Related papers
- Sample based Explanations via Generalized RepresentersChe-Ping Tsai, Chih-Kuan Yeh, Pradeep RavikumarNeurIPS 2023 · 13 citations
- FOCUS: Flexible Optimizable Counterfactual Explanations for Tree EnsemblesAna Lucic, Harrie Oosterhuis, Hinda Haned, Maarten de RijkeAAAI 2022 · 87 citations
- Local Feature Selection without Label or Feature Leakage for Interpretable Machine Learning PredictionsHarrie Oosterhuis, Lijun Lyu, Avishek AnandICML 2024 · 2 citations
- ABLE: Using Adversarial Pairs to Construct Local Models for Explaining Model PredictionsKrishna Khadka, Sunny Shree, Pujan Budhathoki, Yu Lei et al.KDD 2026
- Interpretability Gone Bad: The Role of Bounded Rationality in How Practitioners Understand Machine LearningHarmanpreet Kaur, Matthew R. Conrad, Davis Rule, Cliff Lampe et al.CSCW 2024 · 14 citations
