Label-Free Explainability for Unsupervised Models
Jonathan Crabbé, Mihaela van der Schaar
摘要
Unsupervised black-box models are challenging to interpret. Indeed, most existing explainability methods require labels to select which component(s) of the black-box's output to interpret. In the absence of labels, black-box outputs often are representation vectors whose components do not correspond to any meaningful quantity. Hence, choosing which component(s) to interpret in a label-free unsupervised/self-supervised setting is an important, yet unsolved problem. To bridge this gap in the literature, we introduce two crucial extensions of post-hoc explanation techniques: (1) label-free feature importance and (2) label-free example importance that respectively highlight influential features and training examples for a black-box to construct representations at inference time. We demonstrate that our extensions can be successfully implemented as simple wrappers around many existing feature and example importance methods. We illustrate the utility of our label-free explainability paradigm through a qualitative and quantitative comparison of representation spaces learned by various autoencoders trained on distinct unsupervised tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 被引用 88 次
- Evaluating the Robustness of Interpretability Methods through Explanation Invariance and EquivarianceJonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 被引用 27 次
- Interpreting Unsupervised Anomaly Detection in Security via Rule ExtractionRuoyu Li, Qing Li, Yu Zhang, Dan Zhao 等NeurIPS 2023 · 被引用 18 次
- UNR-Explainer: Counterfactual Explanations for Unsupervised Node Representation Learning ModelsHyunju Kang, Geonhee Han, Hogun ParkICLR 2024 · 被引用 8 次
- Dissect Black Box: Interpreting for Rule-Based Explanations in Unsupervised Anomaly DetectionYu Zhang, Ruoyu Li, Nengwu Wu, Qing Li 等NeurIPS 2024 · 被引用 7 次
它引用的顶会 Paper7
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 被引用 784 次
- Explaining Time Series Predictions with Dynamic MasksJonathan Crabbé, Mihaela van der SchaarICML 2021 · 被引用 115 次
- Explaining Latent Representations with a Corpus of ExamplesJonathan Crabbé, Zhaozhi Qian, Fergus Imrie, Mihaela van der SchaarNeurIPS 2021 · 被引用 48 次
相关 Paper
- FIMAP: Feature Importance by Minimal Adversarial PerturbationMatt Chapman-Rounds, Umang Bhatt, Erik Pazos, Marc-Andre Schulz 等AAAI 2021 · 被引用 14 次
- Feature Importance Explanations for Temporal Black-Box ModelsAkshay Sood, Mark CravenAAAI 2022 · 被引用 24 次
- Accurate and Robust Feature Importance Estimation under Distribution ShiftsJayaraman J. Thiagarajan, Vivek Sivaraman Narayanaswamy, Rushil Anirudh, Peer-Timo Bremer 等AAAI 2021 · 被引用 12 次
- Unsupervised Causal Binary Concepts Discovery with VAE for Black-Box Model ExplanationThien Q. Tran, Kazuto Fukuchi, Youhei Akimoto, Jun SakumaAAAI 2022 · 被引用 11 次
- Generative causal explanations of black-box classifiersMatthew R. O'Shaughnessy, Gregory Canal, Marissa Connor, Christopher Rozell 等NeurIPS 2020 · 被引用 83 次
