Towards Robust Metrics for Concept Representation Evaluation
Mateo Espinosa Zarlenga, Pietro Barbiero, Zohreh Shams, Dmitry Kazhdan, Umang Bhatt, Adrian Weller, Mateja Jamnik
摘要
Recent work on interpretability has focused on concept-based explanations, where deep learning models are explained in terms of high-level units of information, referred to as concepts. Concept learning models, however, have been shown to be prone to encoding impurities in their representations, failing to fully capture meaningful features of their inputs. While concept learning lacks metrics to measure such phenomena, the field of disentanglement learning has explored the related notion of underlying factors of variation in the data, with plenty of metrics to measure the purity of such factors. In this paper, we show that such metrics are not appropriate for concept learning and propose novel metrics for evaluating the purity of concept representations in both approaches. We show the advantage of these metrics over existing ones and demonstrate their utility in evaluating the robustness of concept representations and interventions performed on them. In addition, we show their utility for benchmarking state-of-the-art methods from both families and find that, contrary to common assumptions, supervision alone may not be sufficient for pure concept representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga 等ICML 2023 · 被引用 68 次
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 被引用 52 次
- Interpretable Concept-Based Memory ReasoningDavid Debot, Pietro Barbiero, Francesco Giannini, Gabriele Ciravegna 等NeurIPS 2024 · 被引用 26 次
- Relational Concept Bottleneck ModelsPietro Barbiero, Francesco Giannini, Gabriele Ciravegna, Michelangelo Diligenti 等NeurIPS 2024 · 被引用 21 次
- Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic LensSamuele Bortolotti, Emanuele Marconato, Paolo Morettin, Andrea Passerini 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper4
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf 等ICML 2020 · 被引用 361 次
- Benchmarks, Algorithms, and Metrics for Hierarchical DisentanglementAndrew Slavin Ross, Finale Doshi-VelezICML 2021 · 被引用 15 次
相关 Paper
- Theory and Evaluation Metrics for Learning Disentangled RepresentationsKien Do, Truyen TranICLR 2020 · 被引用 107 次
- On Causally Disentangled RepresentationsAbbavaram Gowtham Reddy, Benin Godfrey L, Vineeth N. BalasubramanianAAAI 2022 · 被引用 30 次
- Where and What? Examining Interpretable Disentangled RepresentationsXinqi Zhu, Chang Xu, Dacheng TaoCVPR 2021
- Concept-based Explanations for Out-of-Distribution DetectorsJihye Choi, Jayaram Raghuram, Ryan Feng, Jiefeng Chen 等ICML 2023 · 被引用 18 次
- Disentanglement Analysis with Partial Information DecompositionSeiya Tokui, Issei SatoICLR 2022 · 被引用 16 次
