Learning to Constrain Policy Optimization with Virtual Trust Region
Hung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen, Kien Do, Sunil Gupta, Svetha Venkatesh
摘要
We introduce a constrained optimization method for policy gradient reinforcement learning, which uses a virtual trust region to regulate each policy update. In addition to using the proximity of one single old policy as the normal trust region, we propose forming a second trust region through another virtual policy representing a wide range of past policies. We then enforce the new policy to stay closer to the virtual policy, which is beneficial if the old policy performs poorly. More importantly, we propose a mechanism to automatically build the virtual policy from a memory of past policies, providing a new capability for dynamically learning appropriate virtual trust regions during the optimization process. Our proposed method, dubbed Memory-Constrained Policy Optimization (MCPO), is examined in diverse environments, including robotic locomotion control, navigation with sparse rewards and Atari games, consistently demonstrating competitive performance against recent on-policy constrained policy gradient methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 被引用 11 次
- Multi-Reference Preference Optimization for Large Language ModelsHung Le, Quan Hung Tran, Dung Nguyen, Kien Do 等AAAI 2025 · 被引用 6 次
它引用的顶会 Paper7
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark 等ICLR 2020 · 被引用 138 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Neural Stored-program MemoryHung Le, Truyen Tran, Svetha VenkateshICLR 2020 · 被引用 38 次
- Deep Conservative Policy IterationNino Vieillard, Olivier Pietquin, Matthieu GeistAAAI 2020 · 被引用 29 次
相关 Paper
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- Supported Trust Region Optimization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu 等ICML 2023 · 被引用 24 次
- Ratio-Variance Regularized Policy OptimizationYu Luo, Shuo Han, Yihan Hu, Lei Lv 等ICML 2026
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song 等NeurIPS 2024 · 被引用 22 次
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen 等ICLR 2025
