Meta-Gradient Reinforcement Learning with an Objective Discovered Online
Zhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, David Silver
Abstract
Deep reinforcement learning includes a broad family of algorithms that parameterise an internal representation, such as a value function or policy, by a deep neural network. Each algorithm optimises its parameters with respect to an objective, such as Q-learning or policy gradient, that defines its semantics. In this work, we propose an algorithm based on meta-gradient descent that discovers its own objective, flexibly parameterised by a deep neural network, solely from interactive experience with its environment. Over time, this allows the agent to learn how to learn increasingly effectively. Furthermore, because the objective is discovered online, it can adapt to changes over time. We demonstrate that the algorithm discovers how to address several important issues in RL, such as bootstrapping, non-stationarity, and off-policy learning. On the Atari Learning Environment, the meta-gradient algorithm adapts over time to learn with greater efficiency, eventually outperforming the median score of a strong actor-critic baseline.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7e2aa1b-d73e-4fd5-89fa-d773ee56b56fCited by top-tier papers25
- Recurrent Model-Free RL Can Be a Strong Baseline for Many POMDPsTianwei Ni, Benjamin Eysenbach, Ruslan SalakhutdinovICML 2022 · 162 citations
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Meta-Learning with Task-Adaptive Loss Function for Few-Shot LearningSungyong Baik, Janghoon Choi, Heewon Kim, Dohee Cho et al.ICCV 2021 · 146 citations
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 96 citations
- Offline Meta-Reinforcement Learning with Online Self-SupervisionVitchyr H. Pong, Ashvin Nair, Laura Smith, Catherine Huang et al.ICML 2022 · 78 citations
Builds on2
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
Related papers
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt et al.ICLR 2022 · 62 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
- Discovering Temporally-Aware Reinforcement Learning AlgorithmsMatthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange et al.ICLR 2024 · 23 citations
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- Discovery of Options via Meta-Learned SubgoalsVivek Veeriah, Tom Zahavy, Matteo Hessel, Zhongwen Xu et al.NeurIPS 2021 · 38 citations
