Rethinking Value Function Learning for Generalization in Reinforcement Learning
Seungyong Moon, JunYeong Lee, Hyun Oh Song
摘要
Our work focuses on training RL agents on multiple visually diverse environments to improve observational generalization performance. In prior methods, policy and value networks are separately optimized using a disjoint network architecture to avoid interference and obtain a more accurate value function. We identify that a value network in the multi-environment setting is more challenging to optimize and prone to memorizing the training data than in the conventional single-environment setting. In addition, we find that appropriate regularization on the value network is necessary to improve both training and test performance. To this end, we propose Delayed-Critic Policy Gradient (DCPG), a policy gradient algorithm that implicitly penalizes value estimates by optimizing the value network less frequently with more training data than the policy network. This can be implemented using a single unified network architecture. Furthermore, we introduce a simple self-supervised task that learns the forward and inverse dynamics of environments using a single discriminator, which can be jointly optimized with the value network. Our proposed algorithms significantly improve observational generalization performance and sample efficiency on the Procgen Benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive LearningSeungyong Moon, Junyoung Yeom, Bumsoo Park, Hyun Oh SongNeurIPS 2023 · 被引用 12 次
- COLA: Towards Efficient Multi-Objective Reinforcement Learning with Conflict Objective Regularization in Latent SpacePengyi Li, Hongyao Tang, Yifu Yuan, Jianye Hao 等NeurIPS 2025 · 被引用 3 次
- PPG Reloaded: An Empirical Study on What Matters in Phasic Policy GradientKaixin Wang, Daquan Zhou, Jiashi Feng, Shie MannorICML 2023 · 被引用 1 次
- Studying the Interplay Between the Actor and Critic Representations in Reinforcement LearningSamuel Garcin, Trevor McInroe, Pablo Samuel Castro, Christopher G. Lucas 等ICLR 2025
- Graph-Theoretic Intrinsic Reward: Guiding RL with Effective ResistanceJatin Chauhan, Shivam Bhardwaj, Aditya Saibewar, Aditya Ramesh 等ICLR 2026
它引用的顶会 Paper22
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 被引用 911 次
- Reinforcement Learning with Augmented DataMichael Laskin, Kimin Lee, Adam Stooke, Lerrel Pinto 等NeurIPS 2020 · 被引用 833 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Improving Sample Efficiency in Model-Free Reinforcement Learning from ImagesDenis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos 等AAAI 2021 · 被引用 506 次
相关 Paper
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 被引用 191 次
- Decoupling Value and Policy for Generalization in Reinforcement LearningRoberta Raileanu, Rob FergusICML 2021 · 被引用 116 次
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska 等ICML 2022 · 被引用 40 次
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov 等NeurIPS 2021 · 被引用 143 次
- Fast Adaptation to New Environments via Policy-Dynamics Value FunctionsRoberta Raileanu, Maxwell Goldstein, Arthur Szlam, Rob FergusICML 2020 · 被引用 27 次
