Policy Regularization with Dataset Constraint for Offline Reinforcement Learning
Yuhang Ran, Yi-Chen Li, Fuxiang Zhang, Zongzhang Zhang, Yang Yu
摘要
We consider the problem of learning the best possible policy from a fixed dataset, known as offline Reinforcement Learning (RL). A common taxonomy of existing offline RL works is policy regularization, which typically constrains the learned policy by distribution or support of the behavior policy. However, distribution and support constraints are overly conservative since they both force the policy to choose similar actions as the behavior policy when considering particular states. It will limit the learned policy's performance, especially when the behavior policy is sub-optimal. In this paper, we find that regularizing the policy towards the nearest state-action pair can be more effective and thus propose Policy Regularization with Dataset Constraint (PRDC). When updating the policy in a given state, PRDC searches the offline dataset for the nearest state-action sample and then restricts the policy with the action of this sample. Unlike previous works, PRDC can guide the policy with better behaviors from the dataset, allowing it to choose actions that do not appear in the dataset along with the given state. It is a softer constraint but still keeps enough conservatism from out-of-distribution actions. Empirical evidence and theoretical analysis show that PRDC can alleviate offline RL's fundamentally challenging value overestimation issue with a bounded performance gap. Moreover, on a set of locomotion and navigation tasks, PRDC achieves state-of-the-art performance compared with existing methods. Code is available at https: //github.com/LAMDA-RL/PRDC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- ACT: Empowering Decision Transformer with Dynamic Programming via Advantage ConditioningChenxiao Gao, Chenyang Wu, Mingjun Cao, Rui Kong 等AAAI 2024 · 被引用 31 次
- Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model DisentanglementZhi Wang, Li Zhang, Wenhao Wu, Yuanheng Zhu 等NeurIPS 2024 · 被引用 31 次
- Doubly Mild Generalization for Offline Reinforcement LearningYixiu Mao, Qi Wang, Yun Qu, Yuhang Jiang 等NeurIPS 2024 · 被引用 30 次
- SEABO: A Simple Search-Based Method for Offline Imitation LearningJiafei Lyu, Xiaoteng Ma, Le Wan, Runze Liu 等ICLR 2024 · 被引用 17 次
- A2PO: Towards Effective Offline Reinforcement Learning from an Advantage-aware PerspectiveYunpeng Qing, Shunyu Liu, Jingyuan Cong, Kaixuan Chen 等NeurIPS 2024 · 被引用 16 次
它引用的顶会 Paper14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- COMBO: Conservative Offline Model-Based Policy OptimizationTianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran 等NeurIPS 2021 · 被引用 549 次
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng 等ICML 2020 · 被引用 539 次
相关 Paper
- State Deviation Correction for Offline Reinforcement LearningHongchang Zhang, Jianzhun Shao, Yuhang Jiang, Shuncheng He 等AAAI 2022 · 被引用 18 次
- Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement LearningTenglong Liu, Yang Li, Yixing Lan, Hao Gao 等ICML 2024 · 被引用 15 次
- Supported Policy Optimization for Offline Reinforcement LearningJialong Wu, Haixu Wu, Zihan Qiu, Jianmin Wang 等NeurIPS 2022 · 被引用 113 次
- Weighted Policy Constraints for Offline Reinforcement LearningZhiyong Peng, Changlin Han, Yadong Liu, Zongtan ZhouAAAI 2023 · 被引用 18 次
- Behaviour Preference Regression for Offline Reinforcement LearningPadmanaba Srinivasan, William KnottenbeltAAAI 2025
