Discovered Policy Optimisation
Chris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz, Christian Schröder de Witt, Jakob N. Foerster
Abstract
Tremendous progress has been made in reinforcement learning (RL) over the past decade. Most of these advancements came through the continual development of new algorithms, which were designed using a combination of mathematical derivations, intuitions, and experimentation. Such an approach of creating algorithms manually is limited by human understanding and ingenuity. In contrast, meta-learning provides a toolkit for automatic machine learning method optimisation, potentially addressing this flaw. However, black-box approaches which attempt to discover RL algorithms with minimal prior structure have thus far not outperformed existing hand-crafted algorithms. Mirror Learning, which includes RL algorithms, such as PPO, offers a potential middle-ground starting point: while every method in this framework comes with theoretical guarantees, components that differentiate them are subject to design. In this paper we explore the Mirror Learning space by meta-learning a "drift" function. We refer to the immediate result as Learnt Policy Optimisation (LPO). By analysing LPO we gain original insights into policy optimisation which we use to formulate a novel, closed-form RL algorithm, Discovered Policy Optimisation (DPO). Our experiments in Brax environments confirm state-of-the-art performance of LPO and DPO, as well as their transfer to unseen settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12a16323-12c5-4b5c-815a-e5ef8fed12e8Cited by top-tier papers36
- Structured State Space Models for In-Context Reinforcement LearningChris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto et al.NeurIPS 2023 · 164 citations
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 96 citations
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan et al.NeurIPS 2024 · 41 citations
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda et al.NeurIPS 2024 · 38 citations
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement LearningJacob Adkins, Michael Bowling, Adam WhiteNeurIPS 2024 · 36 citations
Builds on7
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Meta Learning Backpropagation And Improving ItLouis Kirsch, Jürgen SchmidhuberNeurIPS 2021 · 70 citations
- Mirror Learning: A Unifying Framework of Policy OptimisationJakub Grudzien Kuba, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 41 citations
Related papers
- Introducing Symmetries to Black Box Meta Reinforcement LearningLouis Kirsch, Sebastian Flennerhag, Hado van Hasselt, Abram L. Friesen et al.AAAI 2022 · 34 citations
- SYMBOL: Generating Flexible Black-Box Optimizers through Symbolic Equation LearningJiacheng Chen, Zeyuan Ma, Hongshu Guo, Yining Ma et al.ICLR 2024 · 27 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
- Model-Based Offline Meta-Reinforcement Learning with RegularizationSen Lin, Jialin Wan, Tengyu Xu, Yingbin Liang et al.ICLR 2022 · 20 citations
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 111 citations
