Parameter-Based Value Functions
Francesco Faccio, Louis Kirsch, Jürgen Schmidhuber
摘要
Traditional off-policy actor-critic Reinforcement Learning (RL) algorithms learn value functions of a single target policy. However, when value functions are updated to track the learned policy, they forget potentially useful information about old policies. We introduce a class of value functions called Parameter-based Value Functions (PVFs) whose inputs include the policy parameters. They can generalize across different policies. PVFs can evaluate the performance of any policy given a state, a state-action pair, or a distribution over the RL agent's initial states. First we show how PVFs yield novel off-policy policy gradient theorems. Then we derive off-policy actor-critic algorithms based on PVFs trained by Monte Carlo or Temporal Difference methods. We show how learned PVFs can zero-shot learn new policies that outperform any policy seen during training. Finally our algorithms are evaluated on a selection of discrete and continuous control tasks using shallow policies and deep neural networks. Their performance is comparable to the one of state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 被引用 67 次
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller 等NeurIPS 2021 · 被引用 64 次
- What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorHongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen 等AAAI 2022 · 被引用 20 次
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng 等ICLR 2023 · 被引用 16 次
- Learning to Identify Critical States for Reinforcement Learning from VideosHaozhe Liu, Mingchen Zhuge, Bing Li, Yuhui Wang 等ICCV 2023 · 被引用 14 次
相关 Paper
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo 等ICML 2025
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
- Hypernetworks for Zero-Shot Transfer in Reinforcement LearningSahand Rezaei-Shoshtari, Charlotte Morissette, François Robert Hogan, Gregory Dudek 等AAAI 2023 · 被引用 23 次
- A Unifying Framework of Off-Policy General Value Function EvaluationTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangNeurIPS 2022 · 被引用 2 次
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real 等ICLR 2021 · 被引用 19 次
