Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
Tian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-Zhu
摘要
Recent advances in language models have demonstrated their capability to solve mathematical reasoning problems, achieving near-perfect accuracy on grade-school level math benchmarks like GSM8K. In this paper, we formally study how language models solve these problems. We design a series of controlled experiments to address several fundamental questions: (1) Can language models truly develop reasoning skills, or do they simply memorize templates? (2) What is the model's hidden (mental) reasoning process? (3) Do models solve math questions using skills similar to or different from humans? (4) Do models trained on GSM8K-like datasets develop reasoning skills beyond those necessary for solving GSM8K problems? (5) What mental process causes models to make reasoning mistakes? (6) How large or deep must a model be to effectively solve GSM8K-level math questions? Our study uncovers many hidden mechanisms by which language models solve mathematical questions, providing insights that extend beyond current understandings of LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper65
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du 等ICLR 2026 · 被引用 225 次
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous ThoughtHanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao 等NeurIPS 2025 · 被引用 86 次
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsCharlie Zhang, Graham Neubig, Xiang YueICML 2026 · 被引用 58 次
它引用的顶会 Paper6
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
- What Algorithms can Transformers Learn? A Study in Length GeneralizationHattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin 等ICLR 2024 · 被引用 189 次
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling LawsZeyuan Allen-Zhu, Yuanzhi LiICLR 2025 · 被引用 8 次
- Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math ProblemsTian Ye, Zicheng Xu, Yuanzhi Li, Zeyuan Allen-ZhuICLR 2025 · 被引用 2 次
相关 Paper
- GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversQintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong 等ACL 2024 · 被引用 9 次
- GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language ModelsIman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi, Oncel Tuzel 等ICLR 2025
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu 等ICLR 2024 · 被引用 637 次
- Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo 等EMNLP 2023 · 被引用 2 次
- MathScale: Scaling Instruction Tuning for Mathematical ReasoningZhengyang Tang, Xingxing Zhang, Benyou Wang, Furu WeiICML 2024 · 被引用 163 次
