Probing for Arithmetic Errors in Language Models
Yucheng Sun, Alessandro Stolfo, Mrinmaya Sachan
Abstract
We investigate whether internal activations in language models can be used to detect arithmetic errors.Starting with a controlled setting of 3-digit addition, we show that simple probes can accurately decode both the model's predicted output and the correct answer from hidden states, regardless of whether the model's output is correct.Building on this, we train lightweight error detectors that predict model correctness with over 90% accuracy.We then extend our analysis to structured chain-ofthought traces on addition-only GSM8K problems and find that probes trained on simple arithmetic generalize well to this more complex setting, revealing consistent internal representations.Finally, we demonstrate that these probes can guide selective re-prompting of erroneous reasoning steps, improving task accuracy with minimal disruption to correct outputs.Our findings suggest that arithmetic errors can be anticipated from internal activations alone, and that simple probes offer a viable path toward lightweight model self-correction. 1 * Equal contribution. 1 Our code and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- A model of errors in transformersSuvrat Raju, Praneeth Kumar NetrapalliICML 2026 · 1 citation
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice QuestionsYoonah Park, Haesung Pyun, Yohan JoICML 2026 · 1 citation
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI EvaluationYizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi WangICML 2026 · 1 citation
- Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language ModelsJeesu Jung, Sangkeun JungACL 2026
- Reading Between the Tokens: Improving Preference Predictions through Mechanistic ForecastingSarah Ball, Simeon Allmendinger, Frauke Kreuter, Niklas KühlICML 2026
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
Related papers
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 8 citations
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao et al.AAAI 2025 · 19 citations
- Large Language Models Can Self-Correct with Key Condition VerificationZhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan et al.EMNLP 2024 · 4 citations
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math ProblemsTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 2 citations
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningNing Miao, Yee Whye Teh, Tom RainforthICLR 2024 · 195 citations
