Discriminative Feature Attributions: Bridging Post Hoc Explainability and Inherent Interpretability
Usha Bhalla, Suraj Srinivas, Himabindu Lakkaraju
摘要
With the increased deployment of machine learning models in various real-world applications, researchers and practitioners alike have emphasized the need for explanations of model behaviour. To this end, two broad strategies have been outlined in prior literature to explain models. Post hoc explanation methods explain the behaviour of complex black-box models by identifying features critical to model predictions; however, prior work has shown that these explanations may not be faithful, in that they incorrectly attribute high importance to features that are unimportant or non-discriminative for the underlying task. Inherently interpretable models, on the other hand, circumvent these issues by explicitly encoding explanations into model architecture, meaning their explanations are naturally faithful, but they often exhibit poor predictive performance due to their limited expressive power. In this work, we identify a key reason for the lack of faithfulness of feature attributions: the lack of robustness of the underlying black-box models, especially to the erasure of unimportant distractor features in the input. To address this issue, we propose Distractor Erasure Tuning (DiET), a method that adapts black-box models to be robust to distractor erasure, thus providing discriminative and faithful attributions. This strategy naturally combines the ease of use of post hoc explanations with the faithfulness of inherently interpretable models. We perform extensive experiments on semi-synthetic and real-world datasets and show that DiET produces models that (1) closely approximate the original black-box models they are intended to explain, and (2) yield explanations that match approximate ground truths available by construction. Our code is made public at https://github.com/AI4LIFE-GROUP/DiET.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- An Unsupervised Approach to Achieve Supervised-Level Explainability in Healthcare RecordsJoakim Edin, Maria Maistro, Lars Maaløe, Lasse Borgholt 等EMNLP 2024 · 被引用 4 次
- Generative Model Inversion Through the Lens of the Manifold HypothesisXiong Peng, Bo Han, Fengfei Yu, Tongliang Liu 等NeurIPS 2025 · 被引用 3 次
- Finding NEM-U: Explaining unsupervised representation learning through neural network generated explanation masksBjørn Leth Møller, Christian Igel, Kristoffer Knutsen Wickstrøm, Jon Sporring 等ICML 2024 · 被引用 3 次
- Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions ExplainabilityJoakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen, Tuukka Ruotsalo 等ACL 2025
- From Pixels to Perception: Interpretable Predictions via Instance-wise Grouped Feature SelectionMoritz Vandenhirtz, Julia E. VogtICML 2025
它引用的顶会 Paper5
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- Which Explanation Should I Choose? A Function Approximation Perspective to Characterizing Post Hoc ExplanationsTessa Han, Suraj Srinivas, Himabindu LakkarajuNeurIPS 2022 · 被引用 126 次
- Do Input Gradients Highlight Discriminative Features?Harshay Shah, Prateek Jain, Praneeth NetrapalliNeurIPS 2021 · 被引用 74 次
- B-cos Networks: Alignment is All We Need for InterpretabilityMoritz Böhle, Mario Fritz, Bernt SchieleCVPR 2022 · 被引用 62 次
相关 Paper
- Robust and Stable Black Box ExplanationsHimabindu Lakkaraju, Nino Arsov, Osbert BastaniICML 2020 · 被引用 93 次
- An Additive Instance-Wise Approach to Multi-class Model InterpretationVy Vo, Van Nguyen, Trung Le, Quan Hung Tran 等ICLR 2023
- Selective ExplanationsLucas Monteiro Paes, Dennis Wei, Flávio P. CalmonNeurIPS 2024 · 被引用 4 次
- DISCRET: Synthesizing Faithful Explanations For Treatment Effect EstimationYinjun Wu, Mayank Keoliya, Kan Chen, Neelay Velingker 等ICML 2024 · 被引用 3 次
- Rethinking Robustness of Model AttributionsSandesh Kamath, Sankalp Mittal, Amit Deshpande, Vineeth N. BalasubramanianAAAI 2024 · 被引用 2 次
