ReDit: Reward Dithering for Improved LLM Policy Optimization
Chenxing Wei, Jiarui Yu, Ying He, Hande Dong, Yao Shu, Fei Richard Yu
摘要
DeepSeek-R1 has successfully enhanced Large Language Models (LLMs) reasoning capabilities through its rule-based reward system. While it's a "perfect" reward system that effectively mitigates reward hacking, such reward functions are often discrete. Our experimental observations suggest that discrete rewards can lead to gradient anomaly, unstable optimization, and slow convergence. To address this issue, we propose ReDit (Reward Dithering), a method that dithers the discrete reward signal by adding simple random noise. With this perturbed reward, exploratory gradients are continuously provided throughout the learning process, enabling smoother gradient updates and accelerating convergence. The injected noise also introduces stochasticity into flat reward regions, encouraging the model to explore novel policies and escape local optima. Experiments across diverse tasks and different LLMs demonstrate the effectiveness and efficiency of ReDit. On average, ReDit achieves performance comparable to vanilla GRPO with only approximately 10% the training steps, and furthermore, still exhibits a 4% performance improvement over vanilla GRPO when trained for a similar duration. Visualizations confirm significant mitigation of gradient issues with ReDit. Moreover, theoretical analyses are provided to further validate these advantages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SeRL: Self-play Reinforcement Learning for Large Language Models with Limited DataWenkai Fang, Shunyu Liu, Yang Zhou, Kongcheng Zhang 等NeurIPS 2025 · 被引用 53 次
- Scheduling Your LLM Reinforcement Learning with Reasoning TreesHong Wang, Zhezheng Hao, Jian Luo, Chenxing Wei 等ICLR 2026 · 被引用 16 次
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling ModelsJiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo 等ICLR 2026
- PAFT: Prompt-Agnostic Fine-TuningChenxing Wei, Mingwen Ou, Ying He, Yao Shu 等EMNLP 2025
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei 等ICLR 2026 · 被引用 168 次
相关 Paper
- ETR: Entropy Trend Reward for Efficient Chain-of-Thought ReasoningXuan Xiong, Huan Liu, Li Gu, Zhixiang Chi 等ACL 2026 · 被引用 2 次
- LEASH: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning ModelYanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang 等ACL 2026 · 被引用 8 次
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao 等ICLR 2026 · 被引用 8 次
- Inference-Time Reward Hacking in Large Language ModelsHadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju 等NeurIPS 2025 · 被引用 38 次
- Smarter Not Harder: Generative Process Evaluation with Intrinsic-Signal Driving and Ability‑Adaptive Reward ShapingTao He, Rongchuan Mu, Lizi Liao, Yixin Cao 等ICLR 2026
