Explaining Latent Representations with a Corpus of Examples
Jonathan Crabbé, Zhaozhi Qian, Fergus Imrie, Mihaela van der Schaar
Abstract
Modern machine learning models are complicated. Most of them rely on convoluted latent representations of their input to issue a prediction. To achieve greater transparency than a black-box that connects inputs to predictions, it is necessary to gain a deeper understanding of these latent representations. To that aim, we propose SimplEx: a user-centred method that provides example-based explanations with reference to a freely selected set of examples, called the corpus. SimplEx uses the corpus to improve the user's understanding of the latent space with post-hoc explanations answering two questions: (1) Which corpus examples explain the prediction issued for a given test example? (2) What features of these corpus examples are relevant for the model to relate them to the test example? SimplEx provides an answer by reconstructing the test latent representation as a mixture of corpus latent representations. Further, we propose a novel approach, the Integrated Jacobian, that allows SimplEx to make explicit the contribution of each corpus feature in the mixture. Through experiments on tasks ranging from mortality prediction to image classification, we demonstrate that these decompositions are robust and accurate. With illustrative use cases in medicine, we show that SimplEx empowers the user by highlighting relevant patterns in the corpus that explain model representations. Moreover, we demonstrate how the freedom in choosing the corpus allows the user to have personalized explanations in terms of examples that are meaningful for them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a77a855-18ca-4aff-bdaa-c74632af92f1Cited by top-tier papers13
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
- Visual correspondence-based explanations improve AI robustness and human-AI team accuracyMohammad Reza Taesiri, Giang Nguyen, Anh NguyenNeurIPS 2022 · 57 citations
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 27 citations
- Label-Free Explainability for Unsupervised ModelsJonathan Crabbé, Mihaela van der SchaarICML 2022 · 24 citations
- Spuriosity Rankings: Sorting Data to Measure and Mitigate BiasesMazda Moayeri, Wenxiao Wang, Sahil Singla, Soheil FeiziNeurIPS 2023 · 19 citations
Builds on4
- Algorithmic Transparency via Quantitative Input Influence: Theory and Experiments with Learning SystemsAnupam Datta, Shayak Sen, Yair ZickS&P 2016 · 774 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
- The effectiveness of feature attribution methods and its correlation with automatic evaluation scoresGiang Nguyen, Daeyoung Kim, Anh NguyenNeurIPS 2021 · 128 citations
Related papers
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri et al.EMNLP 2024 · 3 citations
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
- Explanation by Progressive ExaggerationSumedha Singla, Brian Pollack, Junxiang Chen, Kayhan BatmanghelichICLR 2020 · 116 citations
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
