Parameter-Based Value Functions
Francesco Faccio, Louis Kirsch, Jürgen Schmidhuber
Abstract
Traditional off-policy actor-critic Reinforcement Learning (RL) algorithms learn value functions of a single target policy. However, when value functions are updated to track the learned policy, they forget potentially useful information about old policies. We introduce a class of value functions called Parameter-based Value Functions (PVFs) whose inputs include the policy parameters. They can generalize across different policies. PVFs can evaluate the performance of any policy given a state, a state-action pair, or a distribution over the RL agent's initial states. First we show how PVFs yield novel off-policy policy gradient theorems. Then we derive off-policy actor-critic algorithms based on PVFs trained by Monte Carlo or Temporal Difference methods. We show how learned PVFs can zero-shot learn new policies that outperform any policy seen during training. Finally our algorithms are evaluated on a selection of discrete and continuous control tasks using shallow policies and deep neural networks. Their performance is comparable to the one of state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b58d0bbb-a671-41b0-89e7-5fb8558d5b0eCited by top-tier papers9
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 67 citations
- Universal Off-Policy EvaluationYash Chandak, Scott Niekum, Bruno C. da Silva, Erik G. Learned-Miller et al.NeurIPS 2021 · 64 citations
- What about Inputting Policy in Value Function: Policy Representation and Policy-Extended Value Function ApproximatorHongyao Tang, Zhaopeng Meng, Jianye Hao, Chen Chen et al.AAAI 2022 · 20 citations
- ERL-Re: Efficient Evolutionary Reinforcement Learning with Shared State Representation and Individual Policy RepresentationJianye Hao, Pengyi Li, Hongyao Tang, Yan Zheng et al.ICLR 2023 · 16 citations
- Learning to Identify Critical States for Reinforcement Learning from VideosHaozhe Liu, Mingchen Zhuge, Bing Li, Yuhui Wang et al.ICCV 2023 · 14 citations
Related papers
- Wasserstein Policy OptimizationDavid Pfau, Ian Davies, Diana L. Borsa, João Guilherme Madeira Araújo et al.ICML 2025
- Improving Zero-Shot Offline RL via Behavioral Task SamplingNazim Bendib, Nicolas Perrin-Gilbert, Olivier SigaudICML 2026
- Hypernetworks for Zero-Shot Transfer in Reinforcement LearningSahand Rezaei-Shoshtari, Charlotte Morissette, François Robert Hogan, Gregory Dudek et al.AAAI 2023 · 23 citations
- A Unifying Framework of Off-Policy General Value Function EvaluationTengyu Xu, Zhuoran Yang, Zhaoran Wang, Yingbin LiangNeurIPS 2022 · 2 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
