Correcting misinterpretations of additive models
Benedict Clark, Rick Wilming, Hjalmar Schulz, Rustam Zhumagambetov, Danny Panknin, Stefan Haufe
Abstract
Correct model interpretation in high-stakes settings is critical, yet both post-hoc feature attribution methods and so-called intrinsically interpretable models can systematically attribute false-positive importance to non-informative features such as suppressor variables. Specifically, both linear models and their powerful non-linear generalisation such as General Additive Models (GAMs) are susceptible to spurious attributions to suppressors. We present a principled generalisation of activation patterns – originally developed to make linear models interpretable – to additive models, correctly rejecting suppressor effects for non-linear features. This yields PatternGAM, an importance attribution method based on univariate generative surrogate models for the broad family of additive models, and PatternQLR for polynomial models. Empirical evaluations on the XAI-TRIS benchmark with a novel false-negative invariant formulation of the earth mover’s distance accuracy metric demonstrates significant improvements over popular feature attribution methods and the traditional interpretation of additive models. Finally, real-world case studies on the COMPAS and MIMIC-IV datasets provide new insights into the role of specific features by disentangling genuine target-related information from suppression effects that would mislead conventional GAM interpretations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a404e191-abd9-4d8f-a064-7f3eb4992fccBuilds on3
- Neural Additive Models: Interpretable Machine Learning with Neural NetsRishabh Agarwal, Levi Melnick, Nicholas Frosst, Xuezhou Zhang et al.NeurIPS 2021 · 663 citations
- Theoretical Behavior of XAI Methods in the Presence of Suppressor VariablesRick Wilming, Leo Kieslich, Benedict Clark, Stefan HaufeICML 2023 · 17 citations
- Minimizing False-Positive Attributions in Explanations of Non-Linear ModelsAnders Gjølbye, Stefan Haufe, Lars Kai HansenNeurIPS 2025 · 3 citations
Related papers
- Scalable Interpretability via PolynomialsAbhimanyu Dubey, Filip Radenovic, Dhruv MahajanNeurIPS 2022 · 42 citations
- NODE-GAM: Neural Generalized Additive Model for Interpretable Deep LearningChun-Hao Chang, Rich Caruana, Anna GoldenbergICLR 2022 · 114 citations
- How Interpretable and Trustworthy are GAMs?Chun-Hao Chang, Sarah Tan, Benjamin J. Lengerich, Anna Goldenberg et al.KDD 2021 · 60 citations
- Curve Your Enthusiasm: Concurvity Regularization in Differentiable Generalized Additive ModelsJulien Siems, Konstantin Ditschuneit, Winfried Ripken, Alma Lindborg et al.NeurIPS 2023 · 15 citations
- GAM Coach: Towards Interactive and User-centered Algorithmic RecourseZijie J. Wang, Jennifer Wortman Vaughan, Rich Caruana, Duen Horng ChauCHI 2023 · 18 citations
