Defining and Quantifying the Emergence of Sparse Concepts in DNNs
Jie Ren, Mingjie Li, Qirui Chen, Huiqi Deng, Quanshi Zhang
Abstract
This paper aims to illustrate the concept-emerging phenomenon in a trained DNN. Specifically, we find that the inference score of a DNN can be disentangled into the effects of a few interactive concepts. These concepts can be understood as causal patterns in a sparse, symbolic causal graph, which explains the DNN. The faithfulness of using such a causal graph to explain the DNN is theoretically guaranteed, because we prove that the causal graph can well mimic the DNN's outputs on an exponential number of different masked samples. Besides, such a causal graph can be further simplified and re-written as an And-Or graph (AOG), without losing much explanation accuracy. The code is released at https://github.com/sjtu-xai-lab/aog .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df8178c0-0df9-4640-ae62-fa90ff2e74fbCited by top-tier papers20
- Explaining Generalization Power of a DNN Using Interactive ConceptsHuilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu et al.AAAI 2024 · 33 citations
- Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesDongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren et al.NeurIPS 2023 · 28 citations
- Where We Have Arrived in Proving the Emergence of Sparse Interaction Primitives in DNNsQihan Ren, Jiayang Gao, Wen Shen, Quanshi ZhangICLR 2024 · 23 citations
- Towards the Dynamics of a DNN Learning Symbolic InteractionsQihan Ren, Junpeng Zhang, Yang Xu, Yue Xin et al.NeurIPS 2024 · 21 citations
- Defining and extracting generalizable interaction primitives from DNNsLu Chen, Siyu Lou, Benhao Huang, Quanshi ZhangICLR 2024 · 17 citations
Builds on8
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 73 citations
- Interpreting Multivariate Shapley Interactions in DNNsHao Zhang, Yichen Xie, Longjie Zheng, Die Zhang et al.AAAI 2021 · 70 citations
- Interpreting and Boosting Dropout from a Game-Theoretic ViewHao Zhang, Sen Li, Yinchao Ma, Mingjie Li et al.ICLR 2021 · 53 citations
Related papers
- Neuron Dependency Graphs: A Causal Abstraction of Neural NetworksYaojie Hu, Jin TianICML 2022 · 8 citations
- Can We Faithfully Represent Absence States to Compute Shapley Values on a DNN?Jie Ren, Zhanpeng Zhou, Qirui Chen, Quanshi ZhangICLR 2023
- Does a Neural Network Really Encode Symbolic Concepts?Mingjie Li, Quanshi ZhangICML 2023 · 35 citations
- Causal Concept Graph Models: Beyond Causal Opacity in Deep LearningGabriele Dominici, Pietro Barbiero, Mateo Espinosa Zarlenga, Alberto Termine et al.ICLR 2025
- Linear Explanations for Individual NeuronsTuomas P. Oikarinen, Tsui-Wei WengICML 2024 · 18 citations
