The Value-Improvement Path: Towards Better Representations for Reinforcement Learning
Will Dabney, André Barreto, Mark Rowland, Robert Dadashi, John Quan, Marc G. Bellemare, David Silver
摘要
In value-based reinforcement learning (RL), unlike in supervised learning, the agent faces not a single, stationary, approximation problem, but a sequence of value prediction problems. Each time the policy improves, the nature of the problem changes, shifting both the distribution of states and their values. In this paper we take a novel perspective, arguing that the value prediction problems faced by an RL agent should not be addressed in isolation, but rather as a single, holistic, prediction problem. An RL algorithm generates a sequence of policies that, at least approximately, improve towards the optimal policy. We explicitly characterize the associated sequence of value functions and call it the value-improvement path. Our main idea is to approximate the value-improvement path holistically, rather than to solely track the value function of the current policy. Specifically, we discuss the impact that this holistic view of RL has on representation learning. We demonstrate that a representation that spans the past value-improvement path will also provide an accurate value approximation for future policy improvements. We use this insight to better understand existing approaches to auxiliary tasks and to propose new ones. To test our hypothesis empirically, we augmented a standard deep RL agent with an auxiliary task of learning the value-improvement path. In a study of Atari 2600 games, the augmented agent achieved approximately double the mean and median performance of the baseline agent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon 等ICML 2022 · 被引用 269 次
- Pretraining Representations for Data-Efficient Reinforcement LearningMax Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand 等NeurIPS 2021 · 被引用 151 次
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 被引用 151 次
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RLJesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga 等ICML 2024 · 被引用 118 次
相关 Paper
- Value-driven Hindsight ModellingArthur Guez, Fabio Viola, Theophane Weber, Lars Buesing 等NeurIPS 2020 · 被引用 12 次
- Towards a better understanding of representation dynamics under TD-learningYunhao Tang, Rémi MunosICML 2023 · 被引用 3 次
- Learning State Representations from Random Deep Action-conditional PredictionsZeyu Zheng, Vivek Veeriah, Risto Vuorio, Richard L. Lewis 等NeurIPS 2021 · 被引用 6 次
- Boosted Curriculum Reinforcement LearningPascal Klink, Carlo D'Eramo, Jan Peters, Joni PajarinenICLR 2022 · 被引用 11 次
- Predictive auxiliary objectives in deep RL mimic learning in the brainChing Fang, Kim StachenfeldICLR 2024 · 被引用 16 次
