Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
Zeping Yu, Sophia Ananiadou
摘要
We find arithmetic ability resides within a limited number of attention heads, with each head specializing in distinct operations. To delve into the reason, we introduce the Comparative Neuron Analysis (CNA) method, which identifies an internal logic chain consisting of four distinct stages from input to prediction: feature enhancing with shallow FFN neurons, feature transferring by shallow attention layers, feature predicting by arithmetic heads, and prediction enhancing among deep FFN neurons. Moreover, we identify the human-interpretable FFN neurons within both feature-enhancing and feature-predicting stages. These findings lead us to investigate the mechanism of LoRA, revealing that it enhances prediction probabilities by amplifying the coefficient scores of FFN neurons related to predictions. Finally, we apply our method in model pruning for arithmetic tasks and model editing for reducing gender bias. Code is on https://github.com/ zepingyu0512/arithmetic-mechanism .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningGuanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian 等NeurIPS 2025 · 被引用 17 次
- Where Culture Fades: Revealing the Cultural Gap in Text-to-Image GenerationChuancheng Shi, Shangze Li, Shiming Guo, Simiao Xie 等CVPR 2026 · 被引用 14 次
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMsJiandong Shao, Yao Lu, Jianfei YangNeurIPS 2025 · 被引用 8 次
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised LearningXinting Huang, Michael HahnICLR 2026 · 被引用 7 次
- How do Large Language Models Learn In-Context? Query and Key Matrices of In-Context Heads are Two Towers for Metric LearningZeping Yu, Sophia AnaniadouEMNLP 2024 · 被引用 2 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian 等NeurIPS 2020 · 被引用 851 次
相关 Paper
- Arithmetic Without Algorithms: Language Models Solve Math with a Bag of HeuristicsYaniv Nikankin, Anja Reusch, Aaron Mueller, Yonatan BelinkovICLR 2025
- Interpreting Arithmetic Reasoning in Large Language Models using Game-Theoretic InteractionsLeilei Wen, Liwei Zheng, Hongda Li, Lijun Sun 等NeurIPS 2025 · 被引用 1 次
- Interpreting and Improving Large Language Models in Arithmetic CalculationWei Zhang, Chaoqun Wan, Yonggang Zhang, Yiu-ming Cheung 等ICML 2024 · 被引用 47 次
- Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward PassesBryan R. Christ, Zachary Gottesman, Jonathan Kropko, Thomas HartvigsenACL 2025
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 被引用 8 次
