Probing for Arithmetic Errors in Language Models
Yucheng Sun, Alessandro Stolfo, Mrinmaya Sachan
摘要
We investigate whether internal activations in language models can be used to detect arithmetic errors.Starting with a controlled setting of 3-digit addition, we show that simple probes can accurately decode both the model's predicted output and the correct answer from hidden states, regardless of whether the model's output is correct.Building on this, we train lightweight error detectors that predict model correctness with over 90% accuracy.We then extend our analysis to structured chain-ofthought traces on addition-only GSM8K problems and find that probes trained on simple arithmetic generalize well to this more complex setting, revealing consistent internal representations.Finally, we demonstrate that these probes can guide selective re-prompting of erroneous reasoning steps, improving task accuracy with minimal disruption to correct outputs.Our findings suggest that arithmetic errors can be anticipated from internal activations alone, and that simple probes offer a viable path toward lightweight model self-correction. 1 * Equal contribution. 1 Our code and
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A model of errors in transformersSuvrat Raju, Praneeth Kumar NetrapalliICML 2026 · 被引用 1 次
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice QuestionsYoonah Park, Haesung Pyun, Yohan JoICML 2026 · 被引用 1 次
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI EvaluationYizheng Huang, Wenjun Zeng, Aditi Kumaresan, Zi WangICML 2026 · 被引用 1 次
- Tracing Logit Trajectories Across Layer Depth: Dataset-Level Explainability for Language ModelsJeesu Jung, Sangkeun JungACL 2026
- Reading Between the Tokens: Improving Preference Predictions through Mechanistic ForecastingSarah Ball, Simeon Allmendinger, Frauke Kreuter, Niklas KühlICML 2026
它引用的顶会 Paper21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan 等NeurIPS 2023 · 被引用 5,828 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
相关 Paper
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate ItLeonardo Bertolazzi, Philipp Mondorf, Barbara Plank, Raffaella BernardiEMNLP 2025 · 被引用 8 次
- S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical ReasonersYuchen Yan, Jin Jiang, Yang Liu, Yixin Cao 等AAAI 2025 · 被引用 19 次
- Large Language Models Can Self-Correct with Key Condition VerificationZhenyu Wu, Qingkai Zeng, Zhihan Zhang, Zhaoxuan Tan 等EMNLP 2024 · 被引用 4 次
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math ProblemsTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 被引用 2 次
- SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step ReasoningNing Miao, Yee Whye Teh, Tom RainforthICLR 2024 · 被引用 195 次
