Actor-Critic based Improper Reinforcement Learning
Mohammadi Zaki, Avi Mohan, Aditya Gopalan, Shie Mannor
摘要
We consider an improper reinforcement learning setting where a learner is given M base controllers for an unknown Markov decision process, and wishes to combine them optimally to produce a potentially new controller that can outperform each of the base ones. This can be useful in tuning across controllers, learnt possibly in mismatched or simulated environments, to obtain a good controller for a given target environment with relatively few trials. Towards this, we propose two algorithms: (1) a Policy Gradient-based approach; and (2) an algorithm that can switch between a simple Actor-Critic (AC) based scheme and a Natural Actor-Critic (NAC) scheme depending on the available information. Both algorithms operate over a class of improper mixtures of the given controllers. For the first case, we derive convergence rate guarantees assuming access to a gradient oracle. For the AC-based approach we provide convergence rate guarantees to a stationary point in the basic AC case and to a global optimum in the NAC case. Numerical results on (i) the standard control theoretic benchmark of stabilizing an cartpole; and (ii) a constrained queueing task show that our improper policy optimization algorithm can stabilize the system even when the base policies at its disposal are unstable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- Improving Sample Complexity Bounds for (Natural) Actor-Critic AlgorithmsTengyu Xu, Zhe Wang, Yingbin LiangNeurIPS 2020 · 被引用 110 次
- Logarithmic Regret for Learning Linear Quadratic Regulators EfficientlyAsaf B. Cassel, Alon Cohen, Tomer KorenICML 2020 · 被引用 68 次
- Boosting for Control of Dynamical SystemsNaman Agarwal, Nataly Brukhim, Elad Hazan, Zhou LuICML 2020 · 被引用 14 次
相关 Paper
- A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic ApproachSwetha Ganesh, Washim Uddin Mondal, Vaneet AggarwalICML 2025
- Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic AlgorithmYang Xu, Swetha Ganesh, Washim Uddin Mondal, Qinbo Bai 等NeurIPS 2025 · 被引用 8 次
- Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement LearningTianchen Zhou, Hairi, Haibo Yang, Jia Liu 等ICML 2024 · 被引用 4 次
- Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time OraclesBhrij Patel, Wesley A. Suttle, Alec Koppel, Vaneet Aggarwal 等ICML 2024 · 被引用 4 次
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 被引用 21 次
