Discount Factor as a Regularizer in Reinforcement Learning
Ron Amit, Ron Meir, Kamil Ciosek
Abstract
Specifying a Reinforcement Learning (RL) task involves choosing a suitable planning horizon, which is typically modeled by a discount factor. It is known that applying RL algorithms with a lower discount factor can act as a regularizer, improving performance in the limited data regime. Yet the exact nature of this regularizer has not been investigated. In this work, we fill in this gap. For several Temporal-Difference (TD) learning methods, we show an explicit equivalence between using a reduced discount factor and adding an explicit regularization term to the algorithm's loss. Motivated by the equivalence, we empirically study this technique compared to standard L 2 regularization by extensive experiments in discrete and continuous domains, using tabular and functional representations. Our experiments suggest the regularization effectiveness is strongly related to properties of the available data, such as size, distribution, and mixing rate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 222d8dfd-5fdc-4b43-bb7a-ea4715561525Cited by top-tier papers15
- Heuristic-Guided Reinforcement LearningChing-An Cheng, Andrey Kolobov, Adith SwaminathanNeurIPS 2021 · 87 citations
- On-Policy Deep Reinforcement Learning for the Average-Reward CriterionYiming Zhang, Keith W. RossICML 2021 · 59 citations
- Behavior Alignment via Reward Function OptimizationDhawal Gupta, Yash Chandak, Scott M. Jordan, Philip S. Thomas et al.NeurIPS 2023 · 27 citations
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 26 citations
- Rethinking Value Function Learning for Generalization in Reinforcement LearningSeungyong Moon, JunYeong Lee, Hyun Oh SongNeurIPS 2022 · 17 citations
Builds on1
Related papers
- The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement LearningSarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan A. Murphy et al.ICML 2023 · 6 citations
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 2 citations
- Taylor Expansion of Discount FactorsYunhao Tang, Mark Rowland, Rémi Munos, Michal ValkoICML 2021 · 8 citations
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska et al.ICML 2022 · 40 citations
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 2 citations
