Does a Neural Network Really Encode Symbolic Concepts?
Mingjie Li, Quanshi Zhang
Abstract
Recently, a series of studies have tried to extract interactions between input variables modeled by a DNN and define such interactions as concepts encoded by the DNN. However, strictly speaking, there still lacks a solid guarantee whether such interactions indeed represent meaningful concepts. Therefore, in this paper, we examine the trustworthiness of interaction concepts from four perspectives. Extensive empirical studies have verified that a well-trained DNN usually encodes sparse, transferable, and discriminative concepts, which is partially aligned with human intuition. The code is released at https://github.com/sjtu-xai-lab/ interaction-concept .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 006b4be0-bdaa-41c9-8724-2ab4288ab995Cited by top-tier papers8
- Layerwise Change of Knowledge in Neural NetworksXu Cheng, Lei Cheng, Zhaoran Peng, Yang Xu et al.ICML 2024 · 7 citations
- Knowledge Transfer from Interaction LearningYilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu et al.ICCV 2025 · 3 citations
- Evaluating and Explaining Prompt Sensitivity of LLMs Using InteractionsRuiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei et al.ICML 2026 · 1 citation
- A Unified Interpretation of Training-Time Out-Of-Distribution DetectionXu Cheng, Xin Jiang, Zechao LiICCV 2025
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
Builds on13
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 476 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 73 citations
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang et al.AAAI 2021 · 70 citations
Related papers
- Explaining Generalization Power of a DNN Using Interactive ConceptsHuilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu et al.AAAI 2024 · 33 citations
- Defining and Quantifying the Emergence of Sparse Concepts in DNNsJie Ren, Mingjie Li, Qirui Chen, Huiqi Deng et al.CVPR 2023
- Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesDongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren et al.NeurIPS 2023 · 28 citations
- Bayesian Neural Networks Avoid Encoding Complex and Perturbation-Sensitive ConceptsQihan Ren, Huiqi Deng, Yunuo Chen, Siyu Lou et al.ICML 2023 · 13 citations
- Defining and extracting generalizable interaction primitives from DNNsLu Chen, Siyu Lou, Benhao Huang, Quanshi ZhangICLR 2024 · 17 citations
