From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Zishang Jiang, tingyun li, Jinyi Han, Xinyi Wang, Mengyun Qiao, Yizhou Ying, Xiaojun Meng, Jiansheng Wei, Jiaqing Liang, Yanghua Xiao
Abstract
Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longerhorizon interactions. One major bottleneck is distinguishing the contribution of different actions in long-horizon interaction, leading to high optimization variance. To address this, we introduce a novel policy gradient method, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them. We theoretically and empirically show that aggregating semantically similar states and actions in the intent space yields a boundedvariance estimator and improves policy performance stably. Our code is available online 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e78e3f0-be73-42dd-bf77-61d1150e993bBuilds on14
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
- Fine-Tuning Large Vision-Language Models as Decision-Making Agents via Reinforcement LearningSimon Zhai, Hao Bai, Zipeng Lin, Jiayi Pan et al.NeurIPS 2024 · 214 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language ModelsZiniu Li, Tian Xu, Yushun Zhang, Zhihang Lin et al.ICML 2024 · 165 citations
Related papers
- ARIA: Training Language Agents with Intention-driven Reward AggregationRuihan Yang, Yikai Zhang, Aili Chen, Xintao Wang et al.NeurIPS 2025 · 8 citations
- Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language ModelsNianyi Lin, Jiajie Zhang, Lei Hou, Juanzi LiACL 2026 · 8 citations
- Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashionYannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi et al.EMNLP 2024
- RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy OptimizationSiwei Zhang, Yun Xiong, Xi Chen, Zian Jia et al.KDD 2026 · 7 citations
- Retroformer: Retrospective Large Language Agents with Policy Gradient OptimizationWeiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu et al.ICLR 2024 · 124 citations
