Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPs
Siow Meng Low, Akshat Kumar, Scott Sanner
摘要
Recent advances in deep learning have enabled optimization of deep reactive policies (DRPs) for continuous MDP planning by encoding a parametric policy as a deep neural network and exploiting automatic differentiation in an end-to-end model-based gradient descent framework. This approach has proven effective for optimizing DRPs in nonlinear continuous MDPs, but it requires a large number of sampled trajectories to learn effectively and can suffer from high variance in solution quality. In this work, we revisit the overall model-based DRP objective and instead take a minorization-maximization perspective to iteratively optimize the DRP w.r.t. a locally tight lower-bounded objective. This novel formulation of DRP learning as iterative lower bound optimization (ILBO) is particularly appealing because (i) each step is structurally easier to optimize than the overall objective, (ii) it guarantees a monotonically improving objective under certain theoretical conditions, and (iii) it reuses samples between iterations thus lowering sample complexity. Empirical evaluation confirms that ILBO is significantly more sample-efficient than the state-of-the-art DRP planner and consistently produces better solution quality with lower variance. We additionally demonstrate that ILBO generalizes well to new problem instances (i.e., different initial states) without requiring retraining.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Mismatched No More: Joint Model-Policy Optimization for Model-Based RLBenjamin Eysenbach, Alexander Khazatsky, Sergey Levine, Ruslan SalakhutdinovNeurIPS 2022 · 被引用 57 次
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin 等ICLR 2023
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
- The Power of Learned Locally Linear Models for Nonlinear Policy OptimizationDaniel Pfrommer, Max Simchowitz, Tyler Westenbroek, Nikolai Matni 等ICML 2023 · 被引用 4 次
- On the Expressivity of Neural Networks for Deep Reinforcement LearningKefan Dong, Yuping Luo, Tianhe Yu, Chelsea Finn 等ICML 2020 · 被引用 33 次
