Towards Robust Metrics for Concept Representation Evaluation
Mateo Espinosa Zarlenga, Pietro Barbiero, Zohreh Shams, Dmitry Kazhdan, Umang Bhatt, Adrian Weller, Mateja Jamnik
Abstract
Recent work on interpretability has focused on concept-based explanations, where deep learning models are explained in terms of high-level units of information, referred to as concepts. Concept learning models, however, have been shown to be prone to encoding impurities in their representations, failing to fully capture meaningful features of their inputs. While concept learning lacks metrics to measure such phenomena, the field of disentanglement learning has explored the related notion of underlying factors of variation in the data, with plenty of metrics to measure the purity of such factors. In this paper, we show that such metrics are not appropriate for concept learning and propose novel metrics for evaluating the purity of concept representations in both approaches. We show the advantage of these metrics over existing ones and demonstrate their utility in evaluating the robustness of concept representations and interventions performed on them. In addition, we show their utility for benchmarking state-of-the-art methods from both families and find that, contrary to common assumptions, supervision alone may not be sufficient for pure concept representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 207dafd7-4b1b-486b-a0c5-b4a81920777eCited by top-tier papers12
- Interpretable Neural-Symbolic Concept ReasoningPietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga et al.ICML 2023 · 68 citations
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 52 citations
- Interpretable Concept-Based Memory ReasoningDavid Debot, Pietro Barbiero, Francesco Giannini, Gabriele Ciravegna et al.NeurIPS 2024 · 26 citations
- Relational Concept Bottleneck ModelsPietro Barbiero, Francesco Giannini, Gabriele Ciravegna, Michelangelo Diligenti et al.NeurIPS 2024 · 21 citations
- Shortcuts and Identifiability in Concept-based Models from a Neuro-Symbolic LensSamuele Bortolotti, Emanuele Marconato, Paolo Morettin, Andrea Passerini et al.NeurIPS 2025 · 17 citations
Builds on4
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Weakly-Supervised Disentanglement Without CompromisesFrancesco Locatello, Ben Poole, Gunnar Rätsch, Bernhard Schölkopf et al.ICML 2020 · 361 citations
- Benchmarks, Algorithms, and Metrics for Hierarchical DisentanglementAndrew Slavin Ross, Finale Doshi-VelezICML 2021 · 15 citations
Related papers
- Theory and Evaluation Metrics for Learning Disentangled RepresentationsKien Do, Truyen TranICLR 2020 · 107 citations
- On Causally Disentangled RepresentationsAbbavaram Gowtham Reddy, Benin Godfrey L, Vineeth N. BalasubramanianAAAI 2022 · 30 citations
- Where and What? Examining Interpretable Disentangled RepresentationsXinqi Zhu, Chang Xu, Dacheng TaoCVPR 2021
- Concept-based Explanations for Out-of-Distribution DetectorsJihye Choi, Jayaram Raghuram, Ryan Feng, Jiefeng Chen et al.ICML 2023 · 18 citations
- Disentanglement Analysis with Partial Information DecompositionSeiya Tokui, Issei SatoICLR 2022 · 16 citations
