Learning to Constrain Policy Optimization with Virtual Trust Region
Hung Le, Thommen Karimpanal George, Majid Abdolshah, Dung Nguyen, Kien Do, Sunil Gupta, Svetha Venkatesh
Abstract
We introduce a constrained optimization method for policy gradient reinforcement learning, which uses a virtual trust region to regulate each policy update. In addition to using the proximity of one single old policy as the normal trust region, we propose forming a second trust region through another virtual policy representing a wide range of past policies. We then enforce the new policy to stay closer to the virtual policy, which is beneficial if the old policy performs poorly. More importantly, we propose a mechanism to automatically build the virtual policy from a memory of past policies, providing a new capability for dynamically learning appropriate virtual trust regions during the optimization process. Our proposed method, dubbed Memory-Constrained Policy Optimization (MCPO), is examined in diverse environments, including robotic locomotion control, navigation with sparse rewards and Atari games, consistently demonstrating competitive performance against recent on-policy constrained policy gradient methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d655672f-d243-4a82-8e8c-e19f40aa3212Cited by top-tier papers2
- KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample ComplexityGholamali Aminian, Amir Reza Asadi, Idan Shenfeld, Youssef MrouehNeurIPS 2025 · 11 citations
- Multi-Reference Preference Optimization for Large Language ModelsHung Le, Quan Hung Tran, Dung Nguyen, Kien Do et al.AAAI 2025 · 6 citations
Builds on7
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 201 citations
- V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous ControlH. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark et al.ICLR 2020 · 138 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
- Neural Stored-program MemoryHung Le, Truyen Tran, Svetha VenkateshICLR 2020 · 38 citations
- Deep Conservative Policy IterationNino Vieillard, Olivier Pietquin, Matthieu GeistAAAI 2020 · 29 citations
Related papers
- Embedding Safety into RL: A New Take on Trust Region MethodsNikola Milosevic, Johannes Müller, Nico ScherfICML 2025
- Supported Trust Region Optimization for Offline Reinforcement LearningYixiu Mao, Hongchang Zhang, Chen Chen, Yi Xu et al.ICML 2023 · 24 citations
- Ratio-Variance Regularized Policy OptimizationYu Luo, Shuo Han, Yihan Hu, Lei Lv et al.ICML 2026
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song et al.NeurIPS 2024 · 22 citations
- Online Reinforcement Learning in Non-Stationary Context-Driven EnvironmentsPouya Hamadanian, Arash Nasr-Esfahany, Malte Schwarzkopf, Siddhartha Sen et al.ICLR 2025
