Rectify Evaluation Preference: Improving LLMs' Critique on Math Reasoning via Perplexity-aware Reinforcement Learning
Changyuan Tian, Zhicong Lu, Shuang Qian, Nayu Liu, Peiguang Li, Li Jin, Leiyi Hu, Zhizhao Zeng, Sirui Wang, Ke Zeng, Guozhi Cas
Abstract
To improve Multi-step Mathematical Reasoning (MsMR) of Large Language Models (LLMs), it is crucial to obtain scalable supervision from the corpus by automatically critiquing mistakes in the reasoning process of MsMR and rendering a final verdict of the problem-solution. Most existing methods rely on crafting high-quality supervised fine-tuning demonstrations for critiquing capability enhancement and pay little attention to delving into the underlying reason for the poor critiquing performance of LLMs. In this paper, we orthogonally quantify and investigate the potential reason — imbalanced evaluation preference, and conduct a statistical preference analysis. Motivated by the analysis of the reason, a novel perplexity-aware reinforcement learning algorithm is proposed to rectify the evaluation preference, elevating the critiquing capability. Specifically, to probe into LLMs' critiquing characteristics, a One-to-many Problem-Solution (OPS) benchmark is meticulously constructed to quantify the behavior difference of LLMs when evaluating the problem solutions generated by itself and others. Then, to investigate the behavior difference in depth, we conduct a statistical preference analysis oriented on perplexity and find an intriguing phenomenon — "LLMs incline to judge solutions with lower perplexity as correct", which is dubbed as imbalanced evaluation preference. To rectify this preference, we regard perplexity as the baton in the algorithm of Group Relative Policy Optimization, supporting the LLMs to explore trajectories that judge lower perplexity as wrong and higher perplexity as correct. Extensive experimental results on our built OPS and existing available critic benchmarks demonstrate the validity of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe390b02-2d56-4db7-a59f-6181afd2c478Builds on9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu et al.EuroSys 2025 · 61 citations
- SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive SummarizationNayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu et al.ACL 2025 · 7 citations
- HyperMixer: Specializable Hypergraph Channel Mixing for Long-term Multivariate Time Series ForecastingChangyuan Tian, Zhicong Lu, Zequn Zhang, Heming Yang et al.AAAI 2025 · 5 citations
- PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced TuningZhicong Lu, Changyuan Tian, PeiguangLi PeiguangLi, Li Jin et al.ACL 2025 · 4 citations
Related papers
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement LearningZhongwei Wan, Zhihao Dou, Che Liu, Yu Zhang et al.NeurIPS 2025 · 63 citations
- On the Effect of Negative Gradient in Group Relative Deep Reinforcement OptimizationWenlong Deng, Yi Ren, Muchen Li, Danica J. Sutherland et al.NeurIPS 2025 · 36 citations
- Reinforcement Learning for Large Language Models via Group Preference Reward ShapingHuaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao et al.EMNLP 2025
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM ReasoningZiyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng et al.ICLR 2026 · 4 citations
- TPO: Aligning Large Language Models with Multi-branch & Multi-step Preference TreesWeibin Liao, Xu Chu, Yasha WangICLR 2025
