An Analytical Update Rule for General Policy Optimization
Hepeng Li, Nicholas Clavette, Haibo He
摘要
We present an analytical policy update rule that is independent of parametric function approxima-tors. The policy update rule is suitable for optimizing general stochastic policies and has a monotonic improvement guarantee. It is derived from a closed-form solution to trust-region optimization using calculus of variation, following a new theoretical result that tightens existing bounds for policy improvement using trust-region methods. The update rule builds a connection between policy search methods and value function methods. Moreover, off-policy reinforcement learning algorithms can be derived from the update rule since it does not need to compute integration over on-policy states. In addition, the update rule extends immediately to cooperative multi-agent systems when policy updates are performed by one agent at a time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Trust Region Policy Optimisation in Multi-Agent Reinforcement LearningJakub Grudzien Kuba, Ruiqing Chen, Muning Wen, Ying Wen 等ICLR 2022 · 被引用 367 次
- Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement LearningYulai Zhao, Zhuoran Yang, Zhaoran Wang, Jason D. LeeICML 2023 · 被引用 8 次
- Taylor Expansion Policy OptimizationYunhao Tang, Michal Valko, Rémi MunosICML 2020 · 被引用 16 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Absolute Policy Optimization: Enhancing Lower Probability Bound of Performance with High ConfidenceWeiye Zhao, Feihan Li, Yifan Sun, Rui Chen 等ICML 2024 · 被引用 5 次
