Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
Xichen Zhang, Sitong Wu, Yinghao Zhu, Haoru Tan, Shaozuo Yu, Ziyi He, Jiaya Jia
摘要
Reinforcement learning from verifiable rewards has emerged as a powerful technique for enhancing the complex reasoning abilities of Large Language Models (LLMs). However, these methods are fundamentally constrained by the ''learning cliff'' phenomenon: when faced with problems far beyond their current capabilities, models consistently fail, yielding a persistent zero-reward signal. In policy optimization algorithms like GRPO, this collapses the advantage calculation to zero, rendering these difficult problems invisible to the learning gradient and stalling progress. To overcome this, we introduce Scaf-GRPO (Scaffolded Group Relative Policy Optimization), a progressive training framework that strategically provides minimal guidance only when a model's independent learning has plateaued. The framework first diagnoses learning stagnation and then intervenes by injecting tiered in-prompt hints, ranging from abstract concepts to concrete steps, enabling the model to construct a valid solution by itself. Extensive experiments on challenging mathematics benchmarks demonstrate Scaf-GRPO's effectiveness, boosting the pass@1 score of the Qwen2.5-Math-7B model on the AIME24 benchmark by a relative 44.3% over a vanilla GRPO baseline. This result demonstrates our framework provides a robust and effective methodology for unlocking a model's ability to solve problems previously beyond its reach, a critical step towards extending the frontier of autonomous reasoning in LLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- No More Stale Feedback: Co-Evolving Critics for Open-World Agent LearningZhicong Li, Lingjie Jiang, Yulan Hu, Xingchen Zeng 等ACL 2026 · 被引用 3 次
- Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language ModelsYan Liu, Feng Zhang, Zhanyu Ma, Jun Xu 等ACL 2026 · 被引用 2 次
- Advantage Collapse in Group Relative Policy Optimization: Diagnosis and MitigationXixiang He, Qiyao Sun, Ao Cheng, Xingming Li 等ICML 2026
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPOYiming Ren, Yiran Xu, Zicheng Lin, Chufan Shi 等ICML 2026
- CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing AgentsYihong Tang, Kehai Chen, Liang Yue, Benyou Wang 等ICML 2026
它引用的顶会 Paper13
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng 等NeurIPS 2025 · 被引用 592 次
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
相关 Paper
- Do Not Step Into the Same River Twice: Learning to Reason from Trial and ErrorChenming Tang, Hsiu-Yuan Huang, Weijie Liu, Clive Bai 等ACL 2026 · 被引用 2 次
- Advancing LLM Reasoning with Natural Language and Numerical FeedbackXiaoying Zhang, Yipeng Zhang, Hao Sun, Kaituo Feng 等ICML 2026 · 被引用 79 次
- AG-GRPO: Answer-Guided GRPO for Masked Diffusion Language ModelsJuhyeong Kim, Gyunyeop Kim, Sangwoo KangACL 2026
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng 等ICML 2026 · 被引用 17 次
- HiPO: Self-Hint Policy Optimization for RLVRQiyuan Deng, Kehai Chen, Min Zhang, Zhongwen XuICLR 2026
