Explanation by Progressive Exaggeration
Sumedha Singla, Brian Pollack, Junxiang Chen, Kayhan Batmanghelich
摘要
As machine learning methods see greater adoption and implementation in high stakes applications such as medical image diagnosis, the need for model interpretability and explanation has become more critical. Classical approaches that assess feature importance (e.g. saliency maps) do not explain how and why a particular region of an image is relevant to the prediction. We propose a method that explains the outcome of a classification black-box by gradually exaggerating the semantic effect of a given class. Given a query input to a classifier, our method produces a progressive set of plausible variations of that query, which gradually changes the posterior probability from its original class to its negation. These counter-factually generated samples preserve features unrelated to the classification decision, such that a user can employ our method as a "tuning knob" to traverse a data manifold while crossing the decision boundary. Our method is model agnostic and only requires the output value and gradient of the predictor with respect to its input.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Deep Structural Causal Models for Tractable Counterfactual InferenceNick Pawlowski, Daniel Coelho de Castro, Ben GlockerNeurIPS 2020 · 被引用 353 次
- Explaining in Style: Training a GAN to explain a classifier in StyleSpaceOran Lang, Yossi Gandelsman, Michal Yarom, Yoav Wald 等ICCV 2021 · 被引用 181 次
- On Generating Plausible Counterfactual and Semi-Factual Explanations for Deep LearningEoin M. Kenny, Mark T. KeaneAAAI 2021 · 被引用 122 次
- DECE: Decision Explorer with Counterfactual Explanations for Machine Learning ModelsFurui Cheng, Yao Ming, Huamin QuIEEE VIS 2020 · 被引用 118 次
- Meaningfully debugging model mistakes using conceptual counterfactual explanationsAbubakar Abid, Mert Yüksekgönül, James ZouICML 2022 · 被引用 75 次
相关 Paper
- Designing Counterfactual Generators using Deep Model InversionJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Deepta Rajan, Jason Liang 等NeurIPS 2021 · 被引用 25 次
- Accurate Explanation Model for Image Classifiers using Class Association EmbeddingRuitao Xie, Jingbang Chen, Limai Jiang, Rui Xiao 等ICDE 2024 · 被引用 12 次
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell 等NeurIPS 2020 · 被引用 83 次
- Getting a CLUE: A Method for Explaining Uncertainty EstimatesJavier Antorán, Umang Bhatt, Tameem Adel, Adrian Weller 等ICLR 2021 · 被引用 41 次
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
