Frustratingly Easy Regularization on Representation Can Boost Deep Reinforcement Learning
Qiang He, Huangyuan Su, Jieyu Zhang, Xinwen Hou
摘要
Deep reinforcement learning (DRL) gives the promise that an agent learns good policy from high-dimensional information, whereas representation learning removes irrelevant and redundant information and retains pertinent information. In this work, we demonstrate that the learned representation of the 𝑄-network and its target 𝑄-network should, in theory, satisfy a favorable distinguishable representation property. Specifically, there exists an upper bound on the representation similarity of the value functions of two adjacent time steps in a typical DRL setting. However, through illustrative experiments, we show that the learned DRL agent may violate this property and lead to a suboptimal policy. Therefore, we propose a simple yet effective regularizer called Policy Evaluation with Easy Regularization on Representation (PEER), which aims to maintain the distinguishable representation property via explicit regularization on internal representations. And we provide the convergence rate guarantee of PEER. Implementing PEER requires only one line of code. Our experiments demonstrate that incorporating PEER into DRL can significantly improve performance and sample efficiency. Comprehensive experiments show that PEER achieves state-of-the-art performance on all 4 environments on PyBullet, 9 out of 12 tasks on DM-Control, and 19 out of 26 games on Atari. To the best of our knowledge, PEER is the first work to study the inherent representation property of 𝑄-network and its target. Our code is available at https://sites.google.com/view/peer-cvpr2023/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Adaptive Regularization of Representation Rank as an Implicit Constraint of Bellman EquationQiang He, Tianyi Zhou, Meng Fang, Setareh MaghsudiICLR 2024 · 被引用 10 次
- Keep Various Trajectories: Promoting Exploration of Ensemble Policies in Continuous ControlChao Li, Chen Gong, Qiang He, Xinwen HouNeurIPS 2023 · 被引用 8 次
- Advancing DRL Agents in Commercial Fighting Games: Training, Integration, and Agent-Human AlignmentChen Zhang, Qiang He, Yuan Zhou, Elvis S. Liu 等ICML 2024 · 被引用 7 次
- Stabilizing PPO via Latent-Space Regularization and KDE-Driven ExplorationMeiyu Du, Yuqing Gao, Wei WangICML 2026
它引用的顶会 Paper17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 被引用 1,261 次
相关 Paper
- Policy-Independent Behavioral Metric-Based Representation for Deep Reinforcement LearningWeijian Liao, Zongzhang Zhang, Yang YuAAAI 2023 · 被引用 7 次
- DR3: Value-Based Deep Reinforcement Learning Requires Explicit RegularizationAviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville 等ICLR 2022 · 被引用 85 次
- Learning the Target Network in Function SpaceKavosh Asadi, Yao Liu, Shoham Sabach, Ming Yin 等ICML 2024 · 被引用 3 次
- Reining Generalization in Offline Reinforcement Learning via Representation DistinctionYi Ma, Hongyao Tang, Dong Li, Zhaopeng MengNeurIPS 2023 · 被引用 19 次
- Task-Induced Representation LearningJun Yamada, Karl Pertsch, Anisha Gunjal, Joseph J. LimICLR 2022 · 被引用 15 次
