ReDit: Reward Dithering for Improved LLM Policy Optimization
Chenxing Wei, Jiarui Yu, Ying He, Hande Dong, Yao Shu, Fei Richard Yu
Abstract
DeepSeek-R1 has successfully enhanced Large Language Models (LLMs) reasoning capabilities through its rule-based reward system. While it's a "perfect" reward system that effectively mitigates reward hacking, such reward functions are often discrete. Our experimental observations suggest that discrete rewards can lead to gradient anomaly, unstable optimization, and slow convergence. To address this issue, we propose ReDit (Reward Dithering), a method that dithers the discrete reward signal by adding simple random noise. With this perturbed reward, exploratory gradients are continuously provided throughout the learning process, enabling smoother gradient updates and accelerating convergence. The injected noise also introduces stochasticity into flat reward regions, encouraging the model to explore novel policies and escape local optima. Experiments across diverse tasks and different LLMs demonstrate the effectiveness and efficiency of ReDit. On average, ReDit achieves performance comparable to vanilla GRPO with only approximately 10% the training steps, and furthermore, still exhibits a 4% performance improvement over vanilla GRPO when trained for a similar duration. Visualizations confirm significant mitigation of gradient issues with ReDit. Moreover, theoretical analyses are provided to further validate these advantages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- SeRL: Self-play Reinforcement Learning for Large Language Models with Limited DataWenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang et al.NeurIPS 2025 · 53 citations
- Scheduling Your LLM Reinforcement Learning with Reasoning TreesHong Wang, Zhezheng Hao, Jian Luo, Chenxing Wei et al.ICLR 2026 · 16 citations
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling ModelsJiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo et al.ICLR 2026
- PAFT: Prompt-Agnostic Fine-TuningChenxing Wei, Mingwen Ou, Ying He, Yao Shu et al.EMNLP 2025
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei et al.ICLR 2026 · 168 citations
Related papers
- ETR: Entropy Trend Reward for Efficient Chain-of-Thought ReasoningXuan Xiong, Huan Liu, Li Gu, Zhixiang Chi et al.ACL 2026 · 2 citations
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang et al.ACL 2026 · 8 citations
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao et al.ICLR 2026 · 8 citations
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju et al.NeurIPS 2025 · 38 citations
- Smarter Not Harder: Generative Process Evaluation with Intrinsic-Signal Driving and Ability‑Adaptive Reward ShapingTao He, Rongchuan Mu, Lizi Liao, Yixin Cao et al.ICLR 2026
