Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic Interactions
Leilei Wen, Liwei Zheng, Hongda Li, Lijun Sun, Zhihua Wei, Wen Shen
Abstract
In recent years, large language models (LLMs) have made significant advancements in arithmetic reasoning. However, the internal mechanism of how LLMs solve arithmetic problems remains unclear. In this paper, we propose explaining arithmetic reasoning in LLMs using game-theoretic interactions. Specifically, we disentangle the output score of the LLM into numerous interactions between the input words. We quantify different types of interactions encoded by LLMs during forward propagation to explore the internal mechanism of LLMs for solving arithmetic problems. We find that (1) the internal mechanism of LLMs for solving simple one-operator arithmetic problems is their capability to encode operand-operator interactions and high-order interactions from input samples. Additionally, we find that LLMs with weak one-operator arithmetic capabilities focus more on background interactions.
(2) The internal mechanism of LLMs for solving relatively complex two-operator arithmetic problems is their capability to encode operator interactions and operand interactions from input samples. (3) We explain the task-specific nature of the LoRA method from the perspective of interactions. * Corresponding author. 2 Note that each interaction is equivalently encoded by the entire DNN, rather than by a specific neuron. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
operand ๐ผ plus =-1.29 ๐ ๐ = ๐โ๐ ๐ผ ๐ + background interactions ๐ผ What,plus =1.09 ๐ผ plus,is b =1.07 ๐ผ 2 =-2.08 ๐ผ 2,7 =2.27 ๐ผ 2,plus =1.77 ๐ผ plus,7 =1.18 ๐ผ What,is a =-0.72 ๐ผ What,Answer,is b =-0.29 ๐ผ 2,7,is b =-0.50 ๐ผ 2,plus,7 =-0.70
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 ยท 18,833 citations
- Understanding Deep Networks via Extremal Perturbations and Smooth MasksRuth Fong, Mandela Patrick, Andrea VedaldiICCV 2019 ยท 480 citations
- Llemma: An Open Language Model for MathematicsZhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos et al.ICLR 2024 ยท 433 citations
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 ยท 251 citations
- The Shapley Taylor Interaction IndexMukund Sundararajan, Kedar Dhamdhere, Ashish AgarwalICML 2020 ยท 199 citations
Related papers
- Arithmetic Without Algorithms: Language Models Solve Math with a Bag of HeuristicsYaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan BelinkovICLR 2025
- Interpreting and Improving Large Language Models in Arithmetic CalculationWei Zhang, Chaoqun Wan, Yonggang Zhang, Yiu-ming Cheung et al.ICML 2024 ยท 47 citations
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic ComputationZiling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao et al.EMNLP 2025 ยท 6 citations
- An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMsDaking Rai, Ziyu YaoACL 2024
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisAlessandro Stolfo, Yonatan Belinkov, Mrinmaya SachanEMNLP 2023 ยท 11 citations
