Towards the Dynamics of a DNN Learning Symbolic Interactions
Qihan Ren, Junpeng Zhang, Yang Xu, Yue Xin, Dongrui Liu, Quanshi Zhang
Abstract
This study proves the two-phase dynamics of a deep neural network (DNN) learning interactions. Despite the long disappointing view of the faithfulness of post-hoc explanation of a DNN, a series of theorems have been proven in recent years to show that for a given input sample, a small set of interactions between input variables can be considered as primitive inference patterns that faithfully represent a DNN's detailed inference logic on that sample. Particularly, Zhang et al. have observed that various DNNs all learn interactions of different complexities in two distinct phases, and this two-phase dynamics well explains how a DNN changes from under-fitting to over-fitting. Therefore, in this study, we mathematically prove the two-phase dynamics of interactions, providing a theoretical mechanism for how the generalization power of a DNN changes during the training process. Experiments show that our theory well predicts the real dynamics of interactions on different DNNs trained for various tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c28ce25-6500-4a7f-acac-19d2c871c244Cited by top-tier papers7
- ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMsLandon Butler, Abhineet Agarwal, Justin Singh Kang, Yigit Efe Erginbas et al.NeurIPS 2025 · 19 citations
- Learning to Understand: Identifying Interactions via the Möbius TransformJustin Singh Kang, Yigit Efe Erginbas, Landon Butler, Ramtin Pedarsani et al.NeurIPS 2024 · 17 citations
- Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsLeilei Wen, Liwei Zheng, Hongda Li, Lijun Sun et al.NeurIPS 2025 · 1 citation
- A Unified Approach to Interpreting Self-supervised Pre-training Methods for 3D Point Clouds via InteractionsQiang Li, Jian Ruan, Fanghao Wu, Yuchi Chen et al.CVPR 2025
- A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via InteractionsQingzhuo Wang, Ruiyang Qin, Zhenxin Qin, Wen Shen et al.ICML 2026
Builds on10
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 · 199 citations
- A Unified Approach to Interpreting and Boosting Adversarial TransferabilityXin Wang, Jie Ren, Shuyun Lin, Xiangming Zhu et al.ICLR 2021 · 113 citations
- Discovering and Explaining the Representation Bottleneck of DNNSHuiqi Deng, Qihan Ren, Hao Zhang, Quanshi ZhangICLR 2022 · 73 citations
- Explaining Generalization Power of a DNN Using Interactive ConceptsHuilin Zhou, Hao Zhang, Huiqi Deng, Dongrui Liu et al.AAAI 2024 · 33 citations
- Towards the Difficulty for a Deep Neural Network to Learn Concepts of Different ComplexitiesDongrui Liu, Huiqi Deng, Xu Cheng, Qihan Ren et al.NeurIPS 2023 · 28 citations
Related papers
- Monitoring Primitive Interactions During the Training of DNNsJie Ren, Xinhao Zheng, Jiyu Liu, Andrew Lizarraga et al.AAAI 2025 · 4 citations
- Layerwise Change of Knowledge in Neural NetworksXu Cheng, Lei Cheng, Zhaoran Peng, Yang Xu et al.ICML 2024 · 7 citations
- Defining and extracting generalizable interaction primitives from DNNsLu Chen, Siyu Lou, Benhao Huang, Quanshi ZhangICLR 2024 · 17 citations
- Interpreting and Boosting Dropout from a Game-Theoretic ViewHao Zhang, Sen Li, Yinchao Ma, Mingjie Li et al.ICLR 2021 · 53 citations
- Defining and Quantifying the Emergence of Sparse Concepts in DNNsJie Ren, Mingjie Li, Qirui Chen, Huiqi Deng et al.CVPR 2023
