Leveraging Latent Features for Local Explanations
Ronny Luss, Pin-Yu Chen, Amit Dhurandhar, Prasanna Sattigeri, Yunfeng Zhang, Karthikeyan Shanmugam, Chun-Chen Tu
摘要
As the application of deep neural networks proliferates in numerous areas such as medical imaging, video surveillance, and self driving cars, the need for explaining the decisions of these models has become a hot research topic, both at the global and local level. Locally, most explanation methods have focused on identifying relevance of features, limiting the types of explanations possible. In this paper, we investigate a new direction by leveraging latent features to generate contrastive explanations; predictions are explained not only by highlighting aspects that are in themselves sufficient to justify the classification, but also by new aspects which if added will change the classification. The key contribution of this paper lies in how we add features to rich data in a formal yet humanly interpretable way that leads to meaningful results. Our new definition of "addition" uses latent features to move beyond the limitations of previous explanations and resolve an open question laid out in Dhurandhar, et. al. (2018) , which creates local contrastive explanations but is limited to simple datasets such as grayscale images. The strength of our approach in creating intuitive explanations that are also quantitatively superior to other methods is demonstrated on three diverse image datasets (skin lesions, faces, and fashion apparel). A user study with 200 participants further exemplifies the benefits of contrastive information, which can be viewed as complementary to other state-of-the-art interpretability methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- DISSECT: Disentangled Simultaneous Explanations via Concept TraversalsAsma Ghandeharioun, Been Kim, Chun-Liang Li, Brendan Jou 等ICLR 2022 · 被引用 58 次
- On the Safety of Interpretable Machine Learning: A Maximum Deviation ApproachDennis Wei, Rahul Nair, Amit Dhurandhar, Kush R. Varshney 等NeurIPS 2022 · 被引用 12 次
- Impact of Explanation Techniques and Representations on Users' Comprehension and Confidence in Explainable AIJulien Delaunay, Luis Galárraga, Christine Largouët, Niels van BerkelCSCW 2025 · 被引用 6 次
- Let the CAT out of the bag: Contrastive Attributed explanations for TextSaneem A. Chemmengath, Amar Prakash Azad, Ronny Luss, Amit DhurandharEMNLP 2022 · 被引用 6 次
- Interpretable Mesomorphic Networks for Tabular DataArlind Kadra, Sebastian Pineda-Arango, Josif GrabockaNeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper3
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan 等CHI 2021 · 被引用 663 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- GRACE: Generating Concise and Informative Contrastive Sample to Explain Neural Network Model's PredictionThai Le, Suhang Wang, Dongwon LeeKDD 2020 · 被引用 49 次
相关 Paper
- Contrastive Explanations for Model InterpretabilityAlon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar 等EMNLP 2021 · 被引用 12 次
- "Why Not Other Classes?": Towards Class-Contrastive Back-Propagation ExplanationsYipei Wang, Xiaoqian WangNeurIPS 2022 · 被引用 17 次
- Contrastive Corpus Attribution for Explaining RepresentationsChris Lin, Hugh Chen, Chanwoo Kim, Su-In LeeICLR 2023 · 被引用 2 次
- OCTET: Object-aware Counterfactual ExplanationsMehdi Zemni, Mickaël Chen, Éloi Zablocki, Hédi Ben-Younes 等CVPR 2023
- Towards More Faithful Natural Language Explanation Using Multi-Level Contrastive Learning in VQAChengen Lai, Shengli Song, Shiqi Meng, Jingyang Li 等AAAI 2024 · 被引用 12 次
