TIPS: Turn-level Information-Potential Reward Shaping for Search-Augmented LLMs
Yutao Xie, Nathaniel Thomas, Nicklas Hansen, Yang Fu, Li Erran Li, Xiaolong Wang
摘要
Search-augmented large language models (LLMs) trained with reinforcement learning (RL) achieve strong results on open-domain question answering (QA), but training remains brittle: rewards are sparse, credit assignment across reasoning and tool calls is difficult, and optimization often collapses on long-horizon tasks. We introduce Turn-Level Information Potential Reward Shaping (TIPS), a simple RL framework that assigns dense rewards to each reasoning-tool-call segment based on how much it increases a teacher model's log-likelihood of the correct answer. This potential is computed by a frozen or periodically refreshed copy of the policy, so TIPS only requires checkpoints of the model being trained-no separate reward model, verifier, or human process labels-making it practical for scaling to frontier models. We show that this turn-level information reward is a form of potential-based shaping, preserving the task's optimal policy while providing fine-grained guidance beyond outcome-only supervision. On a searchaugmented QA setting spanning seven in-domain and out-of-domain benchmarks, TIPS consistently outperforms PPO/GRPO baselines and substantially improves training stability; for example, on Qwen-2.5-7B Instruct it improves average Exact Match by 11.8% and F1 by 13.6% over PPO. These results suggest that information-potential shaping is a viable general mechanism for stabilizing longhorizon RL on large, tool-using LLMs. The code base for TIPS is available at https://github.com/ucsd-wang-lab-lm/tips .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard 等ICML 2024 · 被引用 598 次
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang 等EMNLP 2023 · 被引用 549 次
相关 Paper
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 被引用 74 次
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan 等ICLR 2026 · 被引用 37 次
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan 等ACL 2026 · 被引用 5 次
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 被引用 484 次
- Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationJunzhe Wang, Zhiheng Xi, Yajie Yang, Hao Luo 等ACL 2026 · 被引用 4 次
