Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods
Chris Nota, Philip S. Thomas, Bruno C. da Silva
摘要
Hindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution over hidden components of the state in previous timesteps and can be used to construct novel unbiased baselines for policy gradient methods. Importantly, we prove that these baselines reduce (and never increase) the variance of policy gradient estimators compared to traditional state value functions. While the posterior value function is motivated by partial observability, we extend these results to arbitrary stochastic MDPs by showing that hindsight-capable agents can model stochasticity in the environment as a special case of partial observability. Finally, we introduce a pair of methods for learning posterior value functions and prove their convergence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor 等ICML 2021 · 被引用 70 次
- Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysisAlexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw 等NeurIPS 2023 · 被引用 16 次
- Policy Gradients Incorporating the FutureDavid Venuto, Elaine Lau, Doina Precup, Ofir NachumICLR 2022 · 被引用 9 次
- Curiosity in Hindsight: Intrinsic Exploration in Stochastic EnvironmentsDaniel Jarrett, Corentin Tallec, Florent Altché, Thomas Mesnard 等ICML 2023 · 被引用 3 次
- Learning to Explore in POMDPs with Informational RewardsAnnie Xie, Logan M. Bhamidipaty, Evan Zheran Liu, Joey Hong 等ICML 2024 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang 等ICML 2023 · 被引用 3 次
- Beyond Variance Reduction: Understanding the True Impact of Baselines on Policy OptimizationWesley Chung, Valentin Thomas, Marlos C. Machado, Nicolas Le RouxICML 2021 · 被引用 35 次
- From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent TrainingZishang Jiang, tingyun li, Jinyi Han, Xinyi Wang 等ICML 2026
- A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature PredictionsAnthony GX-Chen, Veronica Chelu, Blake A. Richards, Joelle PineauAAAI 2022 · 被引用 1 次
- The Role of Baselines in Policy Gradient OptimizationJincheng Mei, Wesley Chung, Valentin Thomas, Bo Dai 等NeurIPS 2022 · 被引用 34 次
