On the Safety of Interpretable Machine Learning: A Maximum Deviation Approach
Dennis Wei, Rahul Nair, Amit Dhurandhar, Kush R. Varshney, Elizabeth Daly, Moninder Singh
Abstract
Interpretable and explainable machine learning has seen a recent surge of interest. We focus on safety as a key motivation behind the surge and make the relationship between interpretability and safety more quantitative. Toward assessing safety, we introduce the concept of maximum deviation via an optimization problem to find the largest deviation of a supervised learning model from a reference model regarded as safe. We then show how interpretability facilitates this safety assessment. For models including decision trees, generalized linear and additive models, the maximum deviation can be computed exactly and efficiently. For tree ensembles, which are not regarded as interpretable, discrete optimization techniques can still provide informative bounds. For a broader class of piecewise Lipschitz functions, we leverage the multi-armed bandit literature to show that interpretability produces tighter (regret) bounds on the maximum deviation. We present case studies, including one on mortgage approval, to illustrate our methods and the insights about models that may be obtained from deviation maximization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Predictive Multiplicity in Probabilistic ClassificationJamelle Watson-Daniels, David C. Parkes, Berk UstunAAAI 2023 · 58 citations
- Additive Models Explained: A Computational Complexity ApproachShahaf Bassan, Michal Moshkovitz, Guy KatzNeurIPS 2025 · 4 citations
- OC-space: a Unifying Perspective on Verification of Tree EnsemblesTimo Martens, Laurens Devos, Lorenzo Cascioli, Wannes Meert et al.ICML 2026
- What Sketch Explainability Really Means for Downstream Tasks?Hmrishav Bandyopadhyay, Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain et al.CVPR 2024
Builds on8
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang et al.NeurIPS 2021 · 663 citations
- Enabling certification of verification-agnostic networks via memory-efficient semidefinite programmingSumanth Dathathri, Krishnamurthy Dvijotham, Alexey Kurakin, Aditi Raghunathan et al.NeurIPS 2020 · 102 citations
- Optimal Counterfactual Explanations in Tree EnsemblesAxel Parmentier, Thibaut VidalICML 2021 · 66 citations
- Synthesizing Programmatic Policies that Inductively GeneralizeJeevana Priya Inala, Osbert Bastani, Zenna Tavares, Armando Solar-LezamaICLR 2020 · 54 citations
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning AgentsChristian Rupprecht, Cyril Ibrahim, Christopher J. PalICLR 2020 · 36 citations
Related papers
- FOCUS: Flexible Optimizable Counterfactual Explanations for Tree EnsemblesAna Lucic, Harrie Oosterhuis, Hinda Haned, Maarten de RijkeAAAI 2022 · 87 citations
- Personal Insights for Altering Decisions of Tree-based Ensembles over TimeNave Frost, Naama Boer, Daniel Deutch, Tova MiloVLDB 2020 · 7 citations
- A Scalable Two Stage Approach to Computing Optimal Decision SetsAlexey Ignatiev, Edward Lam, Peter J. Stuckey, João Marques-SilvaAAAI 2021 · 17 citations
- NODE-GAM: Neural Generalized Additive Model for Interpretable Deep LearningChun-Hao Chang, Rich Caruana, Anna GoldenbergICLR 2022 · 114 citations
- From Black-box to Causal-box: Towards Building More Interpretable ModelsInwoo Hwang, Yushu Pan, Elias BareinboimNeurIPS 2025 · 3 citations
