Distributional Meta-Gradient Reinforcement Learning
Haiyan Yin, Shuicheng Yan, Zhongwen Xu
Abstract
Meta-gradient reinforcement learning (RL) algorithms have substantially boosted the performance of RL agents by learning an adaptive return. All the existing algorithms adhere to the same reward learning principle, where the adaptive return is simply formulated in the form of expected cumulative rewards, upon which the policy and critic update rules are specified under well-adopted distance metrics. In this paper, we present a novel algorithm that builds on the success of metagradient RL algorithms and effectively improves such algorithms by following a simple recipe, i.e., going beyond the expected return to formulate and learn the return in a more expressive form, value distributions. To this end, we first formulate a distributional return that could effectively capture bootstrapping and discounting behaviors over distributions, to form an informative distributional return target in value update. Then we derive an efficient meta update rule to learn the adaptive distributional return with meta-gradients. For empirical evaluation, we first present an illustrative example on a toy two-color grid-world domain, which validates the benefit of learning distributional return over expectation; then we conduct extensive comparisons on a large-scale RL benchmark Atari 2600, where we confirm that our proposed method with distributional return works seamlessly well with the actor-critic framework and leads to state-of-the-art median human normalized score among meta-gradient RL literature.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dd488ef-c950-4290-927e-80b93037fe5eCited by top-tier papers2
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?Guibin Zhang, Junhao Wang, Junjie Chen, Wangchunshu Zhou et al.ICLR 2026 · 107 citations
- Enhancing Online Reinforcement Learning with Meta-Learned Objective from Offline DataShilong Deng, Zetao Zheng, Hongcai He, Paul Weng et al.AAAI 2025
Builds on7
- Meta-Learning with Adaptive HyperparametersSungyong Baik, Myungsub Choi, Janghoon Choi, Heewon Kim et al.NeurIPS 2020 · 164 citations
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
- What Can Learned Intrinsic Rewards Capture?Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu et al.ICML 2020 · 87 citations
Related papers
- Bootstrapped Meta-LearningSebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt et al.ICLR 2022 · 62 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 120 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 58 citations
