Label-Free Explainability for Unsupervised Models
Jonathan Crabbé, Mihaela van der Schaar
Abstract
Unsupervised black-box models are challenging to interpret. Indeed, most existing explainability methods require labels to select which component(s) of the black-box's output to interpret. In the absence of labels, black-box outputs often are representation vectors whose components do not correspond to any meaningful quantity. Hence, choosing which component(s) to interpret in a label-free unsupervised/self-supervised setting is an important, yet unsolved problem. To bridge this gap in the literature, we introduce two crucial extensions of post-hoc explanation techniques: (1) label-free feature importance and (2) label-free example importance that respectively highlight influential features and training examples for a black-box to construct representations at inference time. We demonstrate that our extensions can be successfully implemented as simple wrappers around many existing feature and example importance methods. We illustrate the utility of our label-free explainability paradigm through a qualitative and quantitative comparison of representation spaces learned by various autoencoders trained on distinct unsupervised tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ddcf8e7-36ca-407e-a153-28eb39890797Cited by top-tier papers10
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 27 citations
- Interpreting Unsupervised Anomaly Detection in Security via Rule ExtractionRuoyu Li, Qing Li, Yu Zhang, Dan Zhao et al.NeurIPS 2023 · 18 citations
- UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning ModelsHyunju Kang, Geonhee Han, Hogun ParkICLR 2024 · 8 citations
- Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly DetectionYu Zhang, Ruoyu Li, Nengwu Wu, Qing Li et al.NeurIPS 2024 · 7 citations
Builds on7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 115 citations
- Explaining Latent Representations with a Corpus of ExamplesJonathan Crabbé, Zhaozhi Qian, Fergus Imrie, Mihaela van der SchaarNeurIPS 2021 · 48 citations
Related papers
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz et al.AAAI 2021 · 14 citations
- Feature Importance Explanations for Temporal Black-Box ModelsAkshay Sood, Mark CravenAAAI 2022 · 24 citations
- Accurate and Robust Feature Importance Estimation under Distribution ShiftsJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer et al.AAAI 2021 · 12 citations
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 11 citations
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell et al.NeurIPS 2020 · 83 citations
