Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
Frederik Pahde, Maximilian Dreyer, Moritz Weckbecker, Leander Weber, Christopher J. Anders, Thomas Wiegand, Wojciech Samek, Sebastian Lapuschkin
Abstract
With a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space. Commonly, CAVs are computed by leveraging linear classifiers optimizing the separability of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction. This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability. To address this, we introduce pattern-based CAVs, solely focussing on concept signals, thereby providing more accurate concept directions. We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts. We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures. 1 . RELATED WORK A variety of approaches has emerged to identify human-understandable concepts in DNNs. Some works consider single neurons as concepts (Olah et al., 2017; Achtibat et al., 2023) , while others focus on identifying interesting subspaces (Vielhaben et al., 2023) or linear directions (Nanda et al., 2023) . We follow the latter approach and encode concepts as linear combinations of neurons, also known as superposition (Elhage et al., 2022) . These directions can be identified through unsupervised activation matrix factorization (Fel et al., 2023) or by the supervised training of CAVs, i.e., vectors pointing from samples without to samples with the concept. In the absence of concept labels, automated concept discovery approaches can further streamline this process (Ghorbani et al., 2019; Zhang et al., 2021) . Various methods leverage CAVs as latent concept representation. For instance, TCAV measures a
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8332e1f-1d7d-423d-a4b1-6b292875d091Cited by top-tier papers3
- Manipulating Feature Visualizations with Gradient SlingshotsDilyara Bareeva, Marina M.-C. Höhne, Alexander Warnecke, Lukas Pirch et al.NeurIPS 2025 · 8 citations
- On the Variability of Concept Activation VectorsJulia Wenkmann, Damien GarreauICML 2026 · 3 citations
- FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural NetworksLaines Schmalwasser, Niklas Penzel, Joachim Denzler, Julia NieblingICML 2025
Builds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- EfficientNetV2: Smaller Models and Faster TrainingMingxing Tan, Quoc V. LeICML 2021 · 4,239 citations
- Diffusion Autoencoders: Toward a Meaningful and Decodable RepresentationKonpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, Supasorn SuwajanakornCVPR 2022 · 276 citations
- Invertible Concept-based Explanations for CNN Models with Non-negative Concept Activation VectorsRuihan Zhang, Prashan Madumal, Tim Miller, Krista A. Ehinger et al.AAAI 2021 · 140 citations
- Concept Activation Regions: A Generalized Framework For Concept-Based ExplanationsJonathan Crabbé, Mihaela van der SchaarNeurIPS 2022 · 88 citations
Related papers
- Concept Distillation: Leveraging Human-Centered Explanations for Model ImprovementAvani Gupta, Saurabh Saini, P. J. NarayananNeurIPS 2023 · 18 citations
- Concept Gradient: Concept-based Interpretation Without Linear AssumptionAndrew Bai, Chih-Kuan Yeh, Neil Y. C. Lin, Pradeep Kumar Ravikumar et al.ICLR 2023 · 5 citations
- GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in InterpretabilityZhenghao He, Sanchit Sinha, Guangzhi Xiong, Aidong ZhangICCV 2025 · 2 citations
- LG-CAV: Train Any Concept Activation Vector with Language GuidanceQihan Huang, Jie Song, Mengqi Xue, Haofei Zhang et al.NeurIPS 2024 · 12 citations
- ConEx: Human-Interpretable Saliency Maps via Concept-Aware AttributionYehonatan Elisha, Oren Barkan, Ziv Haddad, Noam KoenigsteinICML 2026
