TIPS: Turn-level Information-Potential Reward Shaping for Search-Augmented LLMs
Yutao Xie, Nathaniel Thomas, Nicklas Hansen, Yang Fu, Li Erran Li, Xiaolong Wang
Abstract
Search-augmented large language models (LLMs) trained with reinforcement learning (RL) achieve strong results on open-domain question answering (QA), but training remains brittle: rewards are sparse, credit assignment across reasoning and tool calls is difficult, and optimization often collapses on long-horizon tasks. We introduce Turn-Level Information Potential Reward Shaping (TIPS), a simple RL framework that assigns dense rewards to each reasoning-tool-call segment based on how much it increases a teacher model's log-likelihood of the correct answer. This potential is computed by a frozen or periodically refreshed copy of the policy, so TIPS only requires checkpoints of the model being trained-no separate reward model, verifier, or human process labels-making it practical for scaling to frontier models. We show that this turn-level information reward is a form of potential-based shaping, preserving the task's optimal policy while providing fine-grained guidance beyond outcome-only supervision. On a searchaugmented QA setting spanning seven in-domain and out-of-domain benchmarks, TIPS consistently outperforms PPO/GRPO baselines and substantially improves training stability; for example, on Qwen-2.5-7B Instruct it improves average Exact Match by 11.8% and F1 by 13.6% over PPO. These results suggest that information-potential shaping is a viable general mechanism for stabilizing longhorizon RL on large, tool-using LLMs. The code base for TIPS is available at https://github.com/ucsd-wang-lab-lm/tips .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62ab282b-a9a9-4d33-a1eb-d3c585eeec8dCited by top-tier papers1
Ask how each one uses itBuilds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon et al.ICML 2023 · 700 citations
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI FeedbackHarrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard et al.ICML 2024 · 598 citations
- G-Eval: NLG Evaluation using Gpt-4 with Better Human AlignmentYang Liu, Dan Iter, Yichong Xu, Shuohang Wang et al.EMNLP 2023 · 549 citations
Related papers
- Dense Reward for Free in Reinforcement Learning from Human FeedbackAlex James Chan, Hao Sun, Samuel Holt, Mihaela van der SchaarICML 2024 · 74 citations
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search AgentsGuoqing Wang, Sunhao Dai, Guangze Ye, Zeyu Gan et al.ICLR 2026 · 37 citations
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy OptimizationYifeng Ding, Hung Le, Songyang Han, Kangrui Ruan et al.ACL 2026 · 5 citations
- Group-in-Group Policy Optimization for LLM Agent TrainingLang Feng, Zhenghai Xue, Tingcong Liu, Bo AnNeurIPS 2025 · 484 citations
- Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy OptimizationJunzhe Wang, Zhiheng Xi, Yajie Yang, Hao Luo et al.ACL 2026 · 4 citations
