Does a Neural Network Really Encode Symbolic Concepts?
Mingjie Li, Quanshi Zhang
2023年份
35被引次数
8顶会引用
摘要
Recently, a series of studies have tried to extract interactions between input variables modeled by a DNN and define such interactions as concepts encoded by the DNN. However, strictly speaking, there still lacks a solid guarantee whether such interactions indeed represent meaningful concepts. Therefore, in this paper, we examine the trustworthiness of interaction concepts from four perspectives. Extensive empirical studies have verified that a well-trained DNN usually encodes sparse, transferable, and discriminative concepts, which is partially aligned with human intuition. The code is released at https://github.com/sjtu-xai-lab/ interaction-concept .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Layerwise Change of Knowledge in Neural NetworksXu Cheng, Lei Cheng, Zhaoran Peng, Yang Xu 等ICML 2024 · 被引用 7 次
- Knowledge Transfer from Interaction LearningYilin Gao, Kangyi Chen, Zhongxing Peng, Hengjie Lu 等ICCV 2025 · 被引用 3 次
- Evaluating and Explaining Prompt Sensitivity of LLMs Using InteractionsRuiyang Qin, Qingzhuo Wang, Tian Wang, Zhihua Wei 等ICML 2026 · 被引用 1 次
- A Unified Interpretation of Training-Time Out-Of-Distribution DetectionXu Cheng, Xin Jiang, Zechao LiICCV 2025
- Rethinking Shapley Value for Negative Interactions in Non-convex GamesWonjoon Chang, Myeongjin Lee, Jaesik ChoiICLR 2025
它引用的顶会 Paper13
- Understanding Global Feature Contributions With Additive Importance MeasuresIan Covert, Scott M. Lundberg, Su-In LeeNeurIPS 2020 · 被引用 476 次
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 被引用 199 次
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu 等ICLR 2021 · 被引用 113 次
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 被引用 73 次
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang 等AAAI 2021 · 被引用 70 次
相关 Paper
- Explaining Generalization Power of a DNN Using Interactive ConceptsHuilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu 等AAAI 2024 · 被引用 33 次
- Defining and Quantifying the Emergence of Sparse Concepts in DNNsJie Ren, Mingjie Li, Qirui Chen, Huiqi Deng 等CVPR 2023
- Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesDongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren 等NeurIPS 2023 · 被引用 28 次
- Bayesian Neural Networks Avoid Encoding Complex and Perturbation-Sensitive ConceptsQihan Ren, Huiqi Deng, Yunuo Chen, Siyu Lou 等ICML 2023 · 被引用 13 次
- Defining and extracting generalizable interaction primitives from DNNsLu Chen, Siyu Lou, Benhao Huang, Quanshi ZhangICLR 2024 · 被引用 17 次
