SmoothHess: ReLU Network Feature Interactions via Stein's Lemma
Max Torop, Aria Masoomi, Davin Hill, Kivanç Köse, Stratis Ioannidis, Jennifer G. Dy
摘要
Several recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-linear and thus have a zero Hessian almost everywhere. We propose SmoothHess, a method of estimating second-order interactions through Stein's Lemma. In particular, we estimate the Hessian of the network convolved with a Gaussian through an efficient sampling algorithm, requiring only network gradient calls. SmoothHess is applied post-hoc, requires no modifications to the ReLU network architecture, and the extent of smoothing can be controlled explicitly. We provide a non-asymptotic bound on the sample complexity of our estimation procedure. We validate the superior ability of SmoothHess to capture interactions on benchmark datasets and a real-world medical spirometry dataset. Related Work Feature Importance and First-Order Methods: Methods that quantify feature importance fall into two categories: (i) perturbation-based methods (e.g., [53, 66, 17] ), which evaluate the change in model outputs with respect to perturbed inputs, and (ii) gradient-based methods (e.g., [72, 76, 81] ), which leverage the natural interpretation of the gradient as infinitesimally local importance for a given sample. Most relevant to our work are gradient-based approaches. The saliency map, as defined in [72] , is simply the gradient of model output with respect to the input. Several variants are developed to address the shortcomings of the saliency maps. SmoothGrad [76] was developed * https://github.com/MaxTorop/SmoothHess * All ReLU network outputs, internal neurons, and SoftMax probabilities are Lipschitz continuous [29, 26] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Reliable Post hoc Explanations: Modeling Uncertainty in ExplainabilityDylan Slack, Anna Hilgard, Sameer Singh, Himabindu LakkarajuNeurIPS 2021 · 被引用 240 次
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 被引用 199 次
- How does This Interaction Affect Me? Interpretable Attribution for Feature InteractionsMichael Tsang, Sirisha Rambhatla, Yan LiuNeurIPS 2020 · 被引用 109 次
- Towards the Unification and Robustness of Perturbation and Gradient Based ExplanationsSushant Agarwal, Shahin Jabbari, Chirag Agarwal, Sohini Upadhyay 等ICML 2021 · 被引用 71 次
- Feature Interaction Interpretability: A Case for Explaining Ad-Recommendation Systems via Neural Interaction DetectionMichael Tsang, Dehua Cheng, Hanpeng Liu, Xue Feng 等ICLR 2020 · 被引用 71 次
相关 Paper
- Neglected Hessian component explains mysteries in sharpness regularizationYann N. Dauphin, Atish Agarwala, Hossein MobahiNeurIPS 2024 · 被引用 16 次
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
- On the Complexity-Faithfulness Trade-Off of Gradient-Based ExplanationsAmir Mehrpanah, Matteo Gamba, Kevin Smith, Hossein AzizpourICCV 2025 · 被引用 2 次
- H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image ClassifiersAyushi Mehrotra, Dipkamal Bhusal, Michael Clifford, Nidhi RastogiCVPR 2026 · 被引用 1 次
- Explaining Local, Global, And Higher-Order Interactions In Deep LearningSamuel Lerman, Charles Venuto, Henry A. Kautz, Chenliang XuICCV 2021 · 被引用 13 次
