GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
Hongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin, Hao Wang, Yifan Wu, Tao Chen, Zhihang Zheng, Tang, Haihua Yang
Abstract
Reinforcement Learning (RL) is pivotal for enhancing Large Language Model (LLM) reasoning, yet mainstream algorithms such as GRPO and DAPO remain constrained by a coarsegrained credit assignment paradigm, where all tokens within the same response receive the identical reward. In this paper, we propose Dynamic Entropy Weighting, systematically define entropy-based weight ratios Hi,t n k=1 H k,t and similar variants to redistribute rewards and get fine-grained rewards through two new algorithms: Group Token Policy Optimization (GTPO), which assigns an entropy-weighted reward to each token and synthesizes token-specific advantage function to drive the model toward optimal path, and the analogous algorithm Sequence-Level GRPO (GRPO-S), which extends this design to the sequence level and exhibits superior stability in long Chain-of-Thought (CoT) reasoning tasks. Unlike methods using entropy as mere regularization, GTPO and GRPO-S establish a new state-of-the-art on AIME and MATH 500, outperforming prior entropy-guided baselines and validating our weighting mechanism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6018adc4-cddd-452d-b183-babf566581a0Cited by top-tier papers10
- Rethinking Entropy Interventions in RLVR: An Entropy Change PerspectiveZhezheng Hao, Hong Wang, Haoyang Liu, Jian Luo et al.ACL 2026 · 42 citations
- DenseGRPO: From Sparse to Dense Reward for Flow Matching Model AlignmentHaoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang et al.ICLR 2026 · 21 citations
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu et al.ACL 2026 · 14 citations
- Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy PrefixesAmrith Setlur, Zijian Wang, Andrew Cohen, Paria Rashidinejad et al.ICML 2026 · 13 citations
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language ModelsYongding Tao, Tian Wang, Yihong Dong, Huanyu Liu et al.ICLR 2026 · 5 citations
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
Related papers
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via sequence-level likelihoodXingyu Lin, Yilin Wen, Du Su, En Wang et al.ACL 2026
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement LearningWenlong Deng, Yi Ren, Yushu Li, Boying Gong et al.ICLR 2026 · 9 citations
- Heterogeneous Adaptive Policy Optimization: Tailoring Optimization to Every Token's NatureZheng Liu, Mengjie Liu, Siwei Wen, Mengzhang Cai et al.ACL 2026 · 9 citations
- KTAE: A Model-Free Algorithm to Key-Tokens Advantage Estimation in Mathematical ReasoningWei Sun, Wen Yang, Pu Jian, Qianlong Du et al.NeurIPS 2025 · 22 citations
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han et al.ICLR 2026 · 47 citations
