An Analytical Update Rule for General Policy Optimization
Hepeng Li, Nicholas Clavette, Haibo He
Abstract
We present an analytical policy update rule that is independent of parametric function approxima-tors. The policy update rule is suitable for optimizing general stochastic policies and has a monotonic improvement guarantee. It is derived from a closed-form solution to trust-region optimization using calculus of variation, following a new theoretical result that tightens existing bounds for policy improvement using trust-region methods. The update rule builds a connection between policy search methods and value function methods. Moreover, off-policy reinforcement learning algorithms can be derived from the update rule since it does not need to compute integration over on-policy states. In addition, the update rule extends immediately to cooperative multi-agent systems when policy updates are performed by one agent at a time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af725a6b-ae65-4b3a-83fe-a1220e7539ddCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen et al.ICLR 2022 · 367 citations
- Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement LearningYulai Zhao, Zhuoran Yang, Zhaoran Wang, Jason D. LeeICML 2023 · 8 citations
- Taylor Expansion Policy OptimizationYunhao Tang, Michal Valko, Rémi MunosICML 2020 · 16 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceWeiye Zhao, Feihan Li, Yifan Sun, Rui Chen et al.ICML 2024 · 5 citations
