On the Safety of Interpretable Machine Learning: A Maximum Deviation Approach
Dennis Wei, Rahul Nair, Amit Dhurandhar, Kush R. Varshney, Elizabeth Daly, Moninder Singh
摘要
Interpretable and explainable machine learning has seen a recent surge of interest. We focus on safety as a key motivation behind the surge and make the relationship between interpretability and safety more quantitative. Toward assessing safety, we introduce the concept of maximum deviation via an optimization problem to find the largest deviation of a supervised learning model from a reference model regarded as safe. We then show how interpretability facilitates this safety assessment. For models including decision trees, generalized linear and additive models, the maximum deviation can be computed exactly and efficiently. For tree ensembles, which are not regarded as interpretable, discrete optimization techniques can still provide informative bounds. For a broader class of piecewise Lipschitz functions, we leverage the multi-armed bandit literature to show that interpretability produces tighter (regret) bounds on the maximum deviation. We present case studies, including one on mortgage approval, to illustrate our methods and the insights about models that may be obtained from deviation maximization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Predictive Multiplicity in Probabilistic ClassificationJamelle Watson-Daniels, David C. Parkes, Berk UstunAAAI 2023 · 被引用 58 次
- Additive Models Explained: A Computational Complexity ApproachShahaf Bassan, Michal Moshkovitz, Guy KatzNeurIPS 2025 · 被引用 4 次
- OC-space: a Unifying Perspective on Verification of Tree EnsemblesTimo Martens, Laurens Devos, Lorenzo Cascioli, Wannes Meert 等ICML 2026
- What Sketch Explainability Really Means for Downstream Tasks?Hmrishav Bandyopadhyay, Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain 等CVPR 2024
它引用的顶会 Paper8
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang 等NeurIPS 2021 · 被引用 663 次
- Enabling certification of verification-agnostic networks via memory-efficient semidefinite programmingSumanth Dathathri, Krishnamurthy Dvijotham, Alexey Kurakin, Aditi Raghunathan 等NeurIPS 2020 · 被引用 102 次
- Optimal Counterfactual Explanations in Tree EnsemblesAxel Parmentier, Thibaut VidalICML 2021 · 被引用 66 次
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 被引用 54 次
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning AgentsChristian Rupprecht, Cyril Ibrahim, Christopher J. PalICLR 2020 · 被引用 36 次
相关 Paper
- FOCUS: Flexible Optimizable Counterfactual Explanations for Tree EnsemblesAna Lucic, Harrie Oosterhuis, Hinda Haned, Maarten de RijkeAAAI 2022 · 被引用 87 次
- Personal Insights for Altering Decisions of Tree-based Ensembles over TimeNave Frost, Naama Boer, Daniel Deutch, Tova MiloVLDB 2020 · 被引用 7 次
- A Scalable Two Stage Approach to Computing Optimal Decision SetsAlexey Ignatiev, Edward Lam, Peter J. Stuckey, João Marques-SilvaAAAI 2021 · 被引用 17 次
- NODE-GAM: Neural Generalized Additive Model for Interpretable Deep LearningChun-Hao Chang, Rich Caruana, Anna GoldenbergICLR 2022 · 被引用 114 次
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 被引用 3 次
