AlphaMath Almost Zero: Process Supervision without Process
Guoxin Chen, Minpeng Liao, Chengxi Li, Kai Fan
摘要
Although recent advancements in large language models (LLMs) have significantly improved their performance on various tasks, they still face challenges with complex and symbolic multi-step reasoning, particularly in mathematical reasoning. To bolster the mathematical reasoning capabilities of LLMs, most existing efforts concentrate on seeking assistance from either domain experts or GPT-4 for high-quality process-supervised data, which is not only expensive but also labor-intensive. In our study, we propose an innovative framework, AlphaMath, that bypasses the need for process annotations (from humans or GPTs) by leveraging Monte Carlo Tree Search (MCTS). This framework focuses on unleashing the potential of a well-pretrained LLM to autonomously enhance its mathematical reasoning. Specifically, we integrate a value model with the LLM, automatically generating both process supervision and step-level evaluation signals in MCTS. Furthermore, we propose an efficient inference strategy, step-level beam search, where the value model is crafted to assist the policy model (i.e., LLM) in navigating more effective reasoning paths, rather than solely relying on prior probabilities. The experimental results on both in-domain and out-of-domain datasets demonstrate that even without GPT-4 or human-annotated process supervision, our AlphaMath framework achieves comparable or superior results to previous state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper79
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-ImprovementXiyao Wang, Zhengyuan Yang, Chao Feng, Hongjin Lu 等NeurIPS 2025 · 被引用 158 次
- Multi-Agent Collaboration via Evolving OrchestrationYufan Dang, Chen Qian, Xueheng Luo, Jingru Fan 等NeurIPS 2025 · 被引用 118 次
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language ModelsYiran Guo, Lijie Xu, Ji Liu, Dan Ye 等NeurIPS 2025 · 被引用 75 次
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree SearchYuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki 等NeurIPS 2025 · 被引用 67 次
- MindJourney: Test-Time Scaling with World Models for Spatial ReasoningYuncong Yang, Jiageng Liu, Zheyuan Zhang, Siyuan Zhou 等NeurIPS 2025 · 被引用 53 次
它引用的顶会 Paper20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin 等NeurIPS 2024 · 被引用 162 次
- AgentPro: Enhancing LLM Agents with Automated Process SupervisionYuchen Deng, Shichen Fan, Naibo Wang, Xinkui Zhao 等EMNLP 2025 · 被引用 2 次
- What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical ReasoningYiran Ma, Zui Chen, Tianqiao Liu, Mi Tian 等AAAI 2025 · 被引用 22 次
- WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructHaipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao 等ICLR 2025
- Progressive Multimodal Reasoning via Active RetrievalGuanting Dong, Chenghao Zhang, Mengjie Deng, Yutao Zhu 等ACL 2025
