ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated Simulatability
Antonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin, Fanny Jourdan
Abstract
Concept-based explanations work by mapping complex model computations to human-understandable concepts. Evaluating such explanations is very difficult, as it includes not only the quality of the induced space of possible concepts but also how effectively the chosen concepts are communicated to users. Existing evaluation metrics often focus solely on the former, neglecting the latter. We introduce an evaluation framework for measuring concept explanations via automated simulatability: a simulator's ability to predict the explained model's outputs based on the provided explanations. This approach accounts for both the concept space and its interpretation in an end-to-end evaluation. Human studies for simulatability are notoriously difficult to enact, particularly at the scale of a wide, comprehensive empirical evaluation (which is the subject of this work). We propose using large language models (LLMs) as simulators to approximate the evaluation and report various analyses to make such approximations reliable. Our method allows for scalable and consistent evaluation across various models and datasets. We report a comprehensive empirical evaluation using this framework and show that LLMs provide consistent rankings of explanation methods. Code available at https://github.com/AnonymousConSim/ConSim.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description FrameworkLaura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer et al.NeurIPS 2025 · 12 citations
- Directly Optimizing Natural Language Explanations for Behavioral Faithfulness: Simulatability and RecoverabilityAdvaith Malladi, Shashank SrivastavaICML 2026
- Explaining Differences Between Model Pairs in Natural Language through Sample LearningAdvaith Malladi, Rakesh R. Menon, Yuvraj Jain, Shashank SrivastavaEMNLP 2025
- Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision ModelsThomas Fel, Ekdeep Singh Lubana, Jacob S. Prince, Matthew Kowal et al.ICML 2025
Builds on9
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Manipulating and Measuring Model InterpretabilityForough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan et al.CHI 2021 · 663 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 147 citations
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger et al.AAAI 2021 · 140 citations
Related papers
- Do Models Explain Themselves? Counterfactual Simulatability of Natural Language ExplanationsYanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao et al.ICML 2024 · 90 citations
- Evaluating Readability and Faithfulness of Concept-based ExplanationsMeng Li, Haoran Jin, Ruixuan Huang, Zhihao Xu et al.EMNLP 2024 · 1 citation
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?Peter Hase, Mohit BansalACL 2020 · 216 citations
- What LLMs Explain Is Not What They Believe: Evaluating Explanation Sufficiency Under Models' Own Input BeliefsNhi Nguyen, Shauli Ravfogel, Rajesh RanganathICML 2026
- Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsBingsheng Yao, Prithviraj Sen, Lucian Popa, James A. Hendler et al.ACL 2023 · 4 citations
