Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Vaishnavi Shrivastava, Ahmed Hassan Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, Dimitris Papailiopoulos
摘要
Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length-inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many tokens are merely "filler": repetitive, verbose text that makes no real progress. We introduce GFPO (Group Filtered Policy Optimization), which curbs this length explosion by sampling larger groups per problem during training and filtering responses to train on based on two key metrics: (1) response length and (2) token efficiency: reward per token ratio. By sampling more at training-time, we teach models to think less at inference-time. On the Phi-4-reasoning model, GFPO cuts GRPO's length inflation by 46-71% across challenging STEM and coding benchmarks (AIME 24/25, GPQA, Omni-MATH, LiveCodeBench) while maintaining accuracy. Optimizing for reward per token further increases reductions in length inflation to 71-85%. We also propose Adaptive Difficulty GFPO, which dynamically allocates more training resources to harder problems based on real-time difficulty estimates, improving the balance between computational efficiency and accuracy especially on difficult questions. GFPO demonstrates that increased training-time compute directly translates to reduced test-time compute-a simple yet effective trade-off for efficient reasoning. GRPO A i,t = R(q,o i )-meanR(q,o 1 ),...,R(q,o G ) stdR(q,o 1 ),...,R(q,o G ) πθ oi,t|q,oi,<t πθ old oi,t|q,oi,<t A (m) i,t , clip πθ oi,t|q,oi,<t πθ old oi,t|q,oi,<t , 1ε, 1 + ε A (m) i,t
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationShih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao 等ICML 2026 · 被引用 128 次
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng 等ICML 2026 · 被引用 21 次
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao 等ICML 2026 · 被引用 20 次
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual ReasoningQi Song, Honglin Li, Yingchen Yu, Haoyi Zhou 等CVPR 2026 · 被引用 16 次
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft LearningYizhou Zhang, Ning Lv, Teng Wang, Jisheng DangICLR 2026 · 被引用 10 次
它引用的顶会 Paper8
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi 等AAAI 2020 · 被引用 395 次
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu 等ICLR 2026 · 被引用 250 次
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang 等NeurIPS 2024 · 被引用 144 次
相关 Paper
- The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingWeize Chen, Jiarui Yuan, Tailin Jin, Ning Ding 等NeurIPS 2025 · 被引用 13 次
- SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model ReasoningChenzhi Hu, Qinzhe Hu, Yuhang Xu, Junyi Chen 等ICML 2026 · 被引用 2 次
- Fast-Slow Thinking GRPO for Large Vision-Language Model ReasoningWenyi Xiao, Leilei GanNeurIPS 2025 · 被引用 34 次
- DRPO: Efficient Reasoning via Decoupled Reward Policy OptimizationGang Li, Yan Chen, Ming Lin, Tianbao YangICLR 2026 · 被引用 19 次
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou 等ACL 2026 · 被引用 5 次
