The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement Learning
Sarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan A. Murphy, Finale Doshi-Velez
摘要
Discount regularization, using a shorter planning horizon when calculating the optimal policy, is a popular choice to restrict planning to a less complex set of policies when estimating an MDP from sparse or noisy data (Jiang et al., 2015). It is commonly understood that discount regularization functions by de-emphasizing or ignoring delayed effects. In this paper, we reveal an alternate view of discount regularization that exposes unintended consequences. We demonstrate that planning under a lower discount factor produces an identical optimal policy to planning using any prior on the transition matrix that has the same distribution for all states and actions. In fact, it functions like a prior with stronger regularization on state-action pairs with more transition data. This leads to poor performance when the transition matrix is estimated from data sets with uneven amounts of data across state-action pairs. Our equivalence theorem leads to an explicit formula to set regularization parameters locally for individual state-action pairs rather than globally. We demonstrate the failures of discount regularization and how we remedy them using our state-action-specific method across simple empirical examples as well as a medical cancer simulator.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Toward Efficient Exploration by Large Language Model AgentsDilip Arumugam, Thomas L. GriffithsICLR 2026 · 被引用 17 次
- Consistent Aggregation of Objectives with Diverse Time Preferences Requires Non-Markovian RewardsSilviu PitisNeurIPS 2023 · 被引用 13 次
它引用的顶会 Paper7
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 被引用 163 次
- The Value Equivalence Principle for Model-Based Reinforcement LearningChristopher Grimm, André Barreto, Satinder Singh, David SilverNeurIPS 2020 · 被引用 129 次
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 被引用 85 次
- Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential TransitionsOmer Gottesman, Joseph Futoma, Yao Liu, Sonali Parbhoo 等ICML 2020 · 被引用 67 次
- Making Sense of Reinforcement Learning and Probabilistic InferenceBrendan O'Donoghue, Ian Osband, Catalin IonescuICLR 2020 · 被引用 54 次
相关 Paper
- Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy SimplicityHarry Chen, Yiyang Sun, Michal Moshkovitz, Zachery Boner 等ICML 2026
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 被引用 26 次
- Correcting discount-factor mismatch in on-policy policy gradient methodsFengdi Che, Gautham Vasan, A. Rupam MahmoodICML 2023 · 被引用 10 次
- Optimistic Planning by Regularized Dynamic ProgrammingAntoine Moulin, Gergely NeuICML 2023 · 被引用 8 次
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 被引用 2 次
