Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human Capability
Vikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong, Olga Russakovsky
摘要
Concept-based interpretability methods aim to explain a deep neural network model's components and predictions using a pre-defined set of semantic concepts. These methods evaluate a trained model on a new, "probe" dataset and correlate the model's outputs with concepts labeled in that dataset. Despite their popularity, they suffer from limitations that are not well-understood and articulated in the literature. In this work, we identify and analyze three commonly overlooked factors in concept-based explanations. First, we find that the choice of the probe dataset has a profound impact on the generated explanations. Our analysis reveals that different probe datasets lead to very different explanations, suggesting that the generated explanations are not generalizable outside the probe dataset. Second, we find that concepts in the probe dataset are often harder to learn than the target classes they are used to explain, calling into question the correctness of the explanations. We argue that only easily learnable concepts should be used in concept-based explanations. Finally, while existing methods use hundreds or even thousands of concepts, our human studies reveal a much stricter upper bound of 32 concepts or less, beyond which the explanations are much less practically useful. We discuss the implications of our findings and provide suggestions for future development of conceptbased interpretability methods. Code for our analysis and user interface can be found at https://github.com/ princetonvisualai/OverlookedFactors
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Coarse-to-Fine Concept Bottleneck ModelsKonstantinos P. Panousis, Dino Ienco, Diego MarcosNeurIPS 2024 · 被引用 35 次
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh 等NeurIPS 2025 · 被引用 23 次
- ConSim: Measuring Concept-Based Explanations' Effectiveness with Automated SimulatabilityAntonin Poché, Alon Jacovi, Agustin Martin Picard, Victor Boutin 等ACL 2025 · 被引用 8 次
- Post-hoc Part-Prototype NetworksAndong Tan, Fengtao Zhou, Hao ChenICML 2024 · 被引用 7 次
- Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of DecodersJames Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou 等NeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper10
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 · 被引用 480 次
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li 等NeurIPS 2020 · 被引用 390 次
- Debugging Tests for Model ExplanationsJulius Adebayo, Michael Muelly, Ilaria Liccardi, Been KimNeurIPS 2020 · 被引用 209 次
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI InteractionSunnie S. Y. Kim, Elizabeth Anne Watkins, Olga Russakovsky, Ruth Fong 等CHI 2023 · 被引用 178 次
相关 Paper
- ConceptExplainer: Interactive Explanation for Deep Neural Networks from a Concept PerspectiveJinbin Huang, Aditi Mishra, Bum Chul Kwon, Chris BryanIEEE VIS 2022 · 被引用 46 次
- Towards Compositionality in Concept LearningAdam Stein, Aaditya Naik, Yinjun Wu, Mayur Naik 等ICML 2024 · 被引用 11 次
- Explanation-based Data Augmentation for Image ClassificationSandareka Wickramanayake, Wynne Hsu, Mong-Li LeeNeurIPS 2021 · 被引用 24 次
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning ModelsZhenge Zhao, Panpan Xu, Carlos Scheidegger, Liu RenIEEE VIS 2021 · 被引用 59 次
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 被引用 1 次
