On the Role of Discount Factor in Offline Reinforcement Learning
Hao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie Zhang
Abstract
Offline reinforcement learning (RL) enables effective learning from previously collected data without exploration, which shows great promise in real-world applications when exploration is expensive or even infeasible. The discount factor, γ, plays a vital role in improving online RL sample efficiency and estimation accuracy, but the role of the discount factor in offline RL is not well explored. This paper examines two distinct effects of γ in offline RL with theoretical analysis, namely the regularization effect and the pessimism effect. On the one hand, γ is a regulator to trade-off optimality with sample efficiency upon existing offline techniques. On the other hand, lower guidance γ can also be seen as a way of pessimism where we optimize the policy's performance in the worst possible models. We empirically verify the above theoretical observation with tabular MDPs and standard D4RL tasks. The results show that the discount factor plays an essential role in the performance of offline RL algorithms, both under small data regimes upon existing offline methods and in large data regimes without other conservative methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35447f06-04ed-4f29-83fa-2b1e2b6454a8Cited by top-tier papers12
- Revisiting the Minimalist Approach to Offline Reinforcement LearningDenis Tarasov, Vladislav Kurenkov, Alexander Nikulin, Sergey KolesnikovNeurIPS 2023 · 148 citations
- Imitation Learning from Observation with Automatic Discount SchedulingYuyang Liu, Weijun Dong, Yingdong Hu, Chuan Wen et al.ICLR 2024 · 15 citations
- Unsupervised Behavior Extraction via Random Intent PriorsHao Hu, Yiqin Yang, Jianing Ye, Ziqing Mai et al.NeurIPS 2023 · 15 citations
- Flow to Control: Offline Reinforcement Learning with Lossless Primitive DiscoveryYiqin Yang, Hao Hu, Wenzhe Li, Siyuan Li et al.AAAI 2023 · 13 citations
- Improving Offline RL by Blending HeuristicsSinong Geng, Aldo Pacchiano, Andrey Kolobov, Ching-An ChengICLR 2024 · 12 citations
Builds on18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
Related papers
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 85 citations
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 39 citations
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
- Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov GameWei Xiong, Han Zhong, Chengshuai Shi, Cong Shen et al.ICLR 2023 · 2 citations
- RORL: Robust Offline Reinforcement Learning via Conservative SmoothingRui Yang, Chenjia Bai, Xiaoteng Ma, Zhaoran Wang et al.NeurIPS 2022 · 118 citations
