The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning
Sarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan A. Murphy, Finale Doshi-Velez
Abstract
Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estimating an MDP from sparse or noisy data (Jiang et al., 2015). It is commonly understood that discount regularization functions by de-emphasizing or ignoring delayed effects. In this paper, we reveal an alternate view of discount regularization that exposes unintended consequences. We demonstrate that planning under a lower discount factor produces an identical optimal policy to planning using any prior on the transition matrix that has the same distribution for all states and actions. In fact, it functions like a prior with stronger regularization on state-action pairs with more transition data. This leads to poor performance when the transition matrix is estimated from data sets with uneven amounts of data across state-action pairs. Our equivalence theorem leads to an explicit formula to set regularization parameters locally for individual state-action pairs rather than globally. We demonstrate the failures of discount regularization and how we remedy them using our state-action-specific method across simple empirical examples as well as a medical cancer simulator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f79930d-1d79-4c41-ad34-fd2e10a74113Cited by top-tier papers2
- Toward Efficient Exploration by Large Language Model AgentsDilip Arumugam, Thomas L. GriffithsICLR 2026 · 17 citations
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 13 citations
Builds on7
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 163 citations
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 129 citations
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 85 citations
- Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential TransitionsOmer Gottesman, Joseph Futoma, Yao Liu, Sonali Parbhoo et al.ICML 2020 · 67 citations
- Making Sense of Reinforcement Learning and Probabilistic InferenceBrendan O'Donoghue, Ian Osband, Catalin IonescuICLR 2020 · 54 citations
Related papers
- Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy SimplicityHarry Chen, Yiyang Sun, Michal Moshkovitz, Zachery Boner et al.ICML 2026
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 26 citations
- Correcting discount-factor mismatch in on-policy policy gradient methodsFengdi Che, Gautham Vasan, A. Rupam MahmoodICML 2023 · 10 citations
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 8 citations
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 2 citations
