Interpreting and Improving Large Language Models in Arithmetic Calculation
Wei Zhang, Chaoqun Wan, Yonggang Zhang, Yiu-ming Cheung, Xinmei Tian, Xu Shen, Jieping Ye
摘要
Large language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathematical computations. However, even for the simplest arithmetic calculations, the intrinsic mechanisms behind LLMs remain mysterious, making it challenging to ensure reliability. In this work, we delve into uncovering a specific mechanism by which LLMs execute calculations. Through comprehensive experiments, we find that LLMs frequently involve a small fraction (< 5%) of attention heads, which play a pivotal role in focusing on operands and operators during calculation processes. Subsequently, the information from these operands is processed through multi-layer perceptrons (MLPs), progressively leading to the final solution. These pivotal heads/MLPs, though identified on a specific dataset, exhibit transferability across different datasets and even distinct tasks. This insight prompted us to investigate the potential benefits of selectively fine-tuning these essential heads/MLPs to boost the LLMs' computational performance. We empirically find that such precise tuning can yield notable enhancements on mathematical prowess, without compromising the performance on non-mathematical tasks. Our work serves as a preliminary exploration into the arithmetic calculation abilities inherent in LLMs, laying a solid foundation to reveal more intricate mathematical tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Unleashing the Power of Large Language Model for Denoising RecommendationShuyao Wang, Zhi Zheng, Yongduo Sui, Hui XiongWWW 2025 · 被引用 18 次
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMsJiandong Shao, Yao Lu, Jianfei YangNeurIPS 2025 · 被引用 8 次
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 被引用 8 次
- A Layer Selection Approach to Test Time AdaptationSabyasachi Sahoo, Mostafa ElAraby, Jonas Ngnawé, Yann Batiste Pequignot 等AAAI 2025 · 被引用 6 次
- Bigram Subnetworks: Mapping to Next Tokens in Transformer Language ModelsTyler A. Chang, Benjamin BergenNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisAlessandro Stolfo, Yonatan Belinkov, Mrinmaya SachanEMNLP 2023 · 被引用 11 次
- Arithmetic Without Algorithms: Language Models Solve Math with a Bag of HeuristicsYaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan BelinkovICLR 2025
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic ComputationZiling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao 等EMNLP 2025 · 被引用 6 次
- Pre-trained Large Language Models Use Fourier Features to Compute AdditionTianyi Zhou, Deqing Fu, Vatsal Sharan, Robin JiaNeurIPS 2024 · 被引用 48 次
- Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsLeilei Wen, Liwei Zheng, Hongda Li, Lijun Sun 等NeurIPS 2025 · 被引用 1 次
