Concept Gradient: Concept-based Interpretation Without Linear Assumption
Andrew Bai, Chih-Kuan Yeh, Neil Y. C. Lin, Pradeep Kumar Ravikumar, Cho-Jui Hsieh
Abstract
Concept-based interpretations of black-box models are often more intuitive than feature-based counterparts for humans to understand. The most widely adopted approach for concept-based gradient interpretation is Concept Activation Vector (CAV). CAV relies on learning linear relations between some latent representations of a given model and concepts. The premise of meaningful concepts lying in a linear subspace of model layers is usually implicitly assumed but does not hold true in general. In this work we proposed Concept Gradients (CG), which extends concept-based gradient interpretation methods to non-linear concept functions. We showed that for a general (potentially non-linear) concept, we can mathematically measure how a small change of concept affects the model's prediction, which is an extension of gradient-based interpretation to the concept space. We demonstrate empirically that CG outperforms CAV in evaluating concept importance on real world datasets and perform a case study on a medical dataset. The code is available at https://github.com/jybai/concept-gradients .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be77655d-5a5d-4737-9ee7-8830aa8cf785Cited by top-tier papers5
- LG-CAV: Train Any Concept Activation Vector with Language GuidanceQihan Huang, Jie Song, Mengqi Xue, Haofei Zhang et al.NeurIPS 2024 · 12 citations
- On the Variability of Concept Activation VectorsJulia Wenkmann, Damien GarreauICML 2026 · 3 citations
- Nonparametric Identification of Latent ConceptsYujia Zheng, Shaoan Xie, Kun ZhangICML 2025
- Improving Adversarial Robustness of Attribution via Implicit RegularizationAmir Mehrpanah, Matteo Gamba, Hossein AzizpourICML 2026
- Large Language Models are Interpretable LearnersRuochen Wang, Si Si, Felix X. Yu, Dorothea Wiesmann Rothuizen et al.ICLR 2025
Builds on7
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Estimating Training Data Influence by Tracing Gradient DescentGarima Pruthi, Frederick Liu, Satyen Kale, Mukund SundararajanNeurIPS 2020 · 784 citations
- On Completeness-aware Concept-Based Explanations in Deep Neural NetworksChih-Kuan Yeh, Been Kim, Sercan Ömer Arik, Chun-Liang Li et al.NeurIPS 2020 · 390 citations
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
- Evaluations and Methods for Explanation through Robustness AnalysisCheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar et al.ICLR 2021 · 68 citations
Related papers
- FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural NetworksLaines Schmalwasser, Niklas Penzel, Joachim Denzler, Julia NieblingICML 2025
- Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional DivergenceFrederik Pahde, Maximilian Dreyer, Moritz Weckbecker, Leander Weber et al.ICLR 2025
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in InterpretabilityZhenghao He, Sanchit Sinha, Guangzhi Xiong, Aidong ZhangICCV 2025 · 2 citations
- Towards Automating Model Explanations with Certified Robustness GuaranteesMengdi Huai, Jinduo Liu, Chenglin Miao, Liuyi Yao et al.AAAI 2022 · 16 citations
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger et al.AAAI 2021 · 140 citations
