On the Role of Discount Factor in Offline Reinforcement Learning
Hao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie Zhang
摘要
Offline reinforcement learning (RL) enables effective learning from previously collected data without exploration, which shows great promise in real-world applications when exploration is expensive or even infeasible. The discount factor, γ, plays a vital role in improving online RL sample efficiency and estimation accuracy, but the role of the discount factor in offline RL is not well explored. This paper examines two distinct effects of γ in offline RL with theoretical analysis, namely the regularization effect and the pessimism effect. On the one hand, γ is a regulator to trade-off optimality with sample efficiency upon existing offline techniques. On the other hand, lower guidance γ can also be seen as a way of pessimism where we optimize the policy's performance in the worst possible models. We empirically verify the above theoretical observation with tabular MDPs and standard D4RL tasks. The results show that the discount factor plays an essential role in the performance of offline RL algorithms, both under small data regimes upon existing offline methods and in large data regimes without other conservative methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Revisiting the Minimalist Approach to Offline Reinforcement LearningDenis Tarasov, Vladislav Kurenkov, Alexander Nikulin, Sergey KolesnikovNeurIPS 2023 · 被引用 148 次
- Imitation Learning from Observation with Automatic Discount SchedulingYuyang Liu, Weijun Dong, Yingdong Hu, Chuan Wen 等ICLR 2024 · 被引用 15 次
- Unsupervised Behavior Extraction via Random Intent PriorsHao Hu, Yiqin Yang, Jianing Ye, Ziqing Mai 等NeurIPS 2023 · 被引用 15 次
- Flow to Control: Offline Reinforcement Learning with Lossless Primitive DiscoveryYiqin Yang, Hao Hu, Wenzhe Li, Siyuan Li 等AAAI 2023 · 被引用 13 次
- Improving Offline RL by Blending HeuristicsSinong Geng, Aldo Pacchiano, Andrey Kolobov, Ching-An ChengICLR 2024 · 被引用 12 次
它引用的顶会 Paper18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
相关 Paper
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 被引用 85 次
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 被引用 39 次
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
- Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov GameWei Xiong, Han Zhong, Chengshuai Shi, Cong Shen 等ICLR 2023 · 被引用 2 次
- RORL: Robust Offline Reinforcement Learning via Conservative SmoothingRui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang 等NeurIPS 2022 · 被引用 118 次
