Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
Vaishnavi Shrivastava, Ahmed Hassan Awadallah, Vidhisha Balachandran, Shivam Garg, Harkirat Behl, Dimitris Papailiopoulos
Abstract
Large language models trained with reinforcement learning with verifiable rewards tend to trade accuracy for length-inflating response lengths to achieve gains in accuracy. While longer answers may be warranted for harder problems, many tokens are merely "filler": repetitive, verbose text that makes no real progress. We introduce GFPO (Group Filtered Policy Optimization), which curbs this length explosion by sampling larger groups per problem during training and filtering responses to train on based on two key metrics: (1) response length and (2) token efficiency: reward per token ratio. By sampling more at training-time, we teach models to think less at inference-time. On the Phi-4-reasoning model, GFPO cuts GRPO's length inflation by 46-71% across challenging STEM and coding benchmarks (AIME 24/25, GPQA, Omni-MATH, LiveCodeBench) while maintaining accuracy. Optimizing for reward per token further increases reductions in length inflation to 71-85%. We also propose Adaptive Difficulty GFPO, which dynamically allocates more training resources to harder problems based on real-time difficulty estimates, improving the balance between computational efficiency and accuracy especially on difficult questions. GFPO demonstrates that increased training-time compute directly translates to reduced test-time compute-a simple yet effective trade-off for efficient reasoning. GRPO A i,t = R(q,o i )-meanR(q,o 1 ),...,R(q,o G ) stdR(q,o 1 ),...,R(q,o G ) πθ oi,t|q,oi,<t πθ old oi,t|q,oi,<t A (m) i,t , clip πθ oi,t|q,oi,<t πθ old oi,t|q,oi,<t , 1ε, 1 + ε A (m) i,t
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL OptimizationShih-Yang Liu, Xin Dong, Ximing Lu, Shizhe Diao et al.ICML 2026 · 128 citations
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng et al.ICML 2026 · 21 citations
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking TokensWei-Lin Chen, Liqian Peng, Tian Tan, Chao Zhao et al.ICML 2026 · 20 citations
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual ReasoningQi Song, Honglin Li, Yingchen Yu, Haoyi Zhou et al.CVPR 2026 · 16 citations
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft LearningYizhou Zhang, Ning Lv, Teng Wang, Jisheng DangICLR 2026 · 10 citations
Builds on8
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi et al.AAAI 2020 · 395 citations
- Dynamic Early Exit in Reasoning ModelsChenxu Yang, Qingyi Si, Yongjie Duan, Zheliang Zhu et al.ICLR 2026 · 250 citations
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang et al.NeurIPS 2024 · 144 citations
Related papers
- The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingWeize Chen, Jiarui Yuan, Tailin Jin, Ning Ding et al.NeurIPS 2025 · 13 citations
- SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model ReasoningChenzhi Hu, Qinzhe Hu, Yuhang Xu, Junyi Chen et al.ICML 2026 · 2 citations
- Fast-Slow Thinking GRPO for Large Vision-Language Model ReasoningWenyi Xiao, Leilei GanNeurIPS 2025 · 34 citations
- DRPO: Efficient Reasoning via Decoupled Reward Policy OptimizationGang Li, Yan Chen, Ming Lin, Tianbao YangICLR 2026 · 19 citations
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou et al.ACL 2026 · 5 citations
