Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human Capability
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong, Olga Russakovsky
Abstract
Concept-based interpretability methods aim to explain a deep neural network model's components and predictions using a pre-defined set of semantic concepts. These methods evaluate a trained model on a new, "probe" dataset and correlate the model's outputs with concepts labeled in that dataset. Despite their popularity, they suffer from limitations that are not well-understood and articulated in the literature. In this work, we identify and analyze three commonly overlooked factors in concept-based explanations. First, we find that the choice of the probe dataset has a profound impact on the generated explanations. Our analysis reveals that different probe datasets lead to very different explanations, suggesting that the generated explanations are not generalizable outside the probe dataset. Second, we find that concepts in the probe dataset are often harder to learn than the target classes they are used to explain, calling into question the correctness of the explanations. We argue that only easily learnable concepts should be used in concept-based explanations. Finally, while existing methods use hundreds or even thousands of concepts, our human studies reveal a much stricter upper bound of 32 concepts or less, beyond which the explanations are much less practically useful. We discuss the implications of our findings and provide suggestions for future development of conceptbased interpretability methods. Code for our analysis and user interface can be found at https://github.com/ princetonvisualai/OverlookedFactors
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Coarse-to-Fine Concept Bottleneck ModelsKonstantinos P. Panousis, Dino Ienco, Diego MarcosNeurIPS 2024 · 35 citations
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh et al.NeurIPS 2025 · 23 citations
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin et al.ACL 2025 · 8 citations
- Post-hoc Part-Prototype NetworksAndong Tan, Fengtao Zhou, Hao ChenICML 2024 · 7 citations
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of DecodersJames Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou et al.NeurIPS 2025 · 7 citations
Builds on10
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 480 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 209 citations
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong et al.CHI 2023 · 178 citations
Related papers
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 46 citations
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik et al.ICML 2024 · 11 citations
- Explanation-based Data Augmentation for Image ClassificationSandareka Wickramanayake, Wynne Hsu, Mong-Li LeeNeurIPS 2021 · 24 citations
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning ModelsZhenge Zhao, Panpan Xu, Carlos Scheidegger, Liu RenIEEE VIS 2021 · 59 citations
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 1 citation
