Noise as a Natural Regularizer in Markov Decision Processes: Connecting Environmental Stochasticity and Policy Simplicity
Harry Chen, Yiyang Sun, Michal Moshkovitz, Zachery Boner, Lesia Semenova, Cynthia Rudin, Ron Parr
摘要
The planning horizon in a Markov Decision Process (MDP) determines how far into the future an agent reasons. In practice, shorter horizons are commonly associated with policies that exhibit simpler or more interpretable decision-making behavior. In this paper, we establish a formal connection between environmental stochasticity and planning horizon in MDPs. We show that for broad classes of transition noise, solving a noisy MDP can be formally related to solving a noise-free MDP with a shorter effective discount factor, leading to identical optimal policies in some cases and near-optimal ones in others. We further characterize settings in which this correspondence breaks down, clarifying when horizon-based interpretations of noise are not valid. These results, which are supported by both theory and experiments, also give some insight into the common practice of using smaller discount factors for reinforcement learning than those that can be justified by standard modeling interpretations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Discount Factor as a Regularizer in Reinforcement LearningRon Amit, Ron Meir, Kamil CiosekICML 2020 · 被引用 85 次
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 被引用 41 次
- On the Role of Discount Factor in Offline Reinforcement LearningHao Hu, Yiqin Yang, Qianchuan Zhao, Chongjie ZhangICML 2022 · 被引用 26 次
- Using Noise to Infer Aspects of Simplicity Without LearningZachery Boner, Harry Chen, Lesia Semenova, Ronald Parr 等NeurIPS 2024 · 被引用 10 次
- Towards Interpretable Deep Reinforcement Learning with Human-Friendly PrototypesEoin M. Kenny, Mycal Tucker, Julie ShahICLR 2023
相关 Paper
- On Shallow Planning Under Partial ObservabilityRandy Lefebvre, Audrey DurandAAAI 2025 · 被引用 2 次
- The Unintended Consequences of Discount Regularization: Improving Regularization in Certainty Equivalence Reinforcement LearningSarah Rathnam, Sonali Parbhoo, Weiwei Pan, Susan A. Murphy 等ICML 2023 · 被引用 6 次
- Settling the Horizon-Dependence of Sample Complexity in Reinforcement LearningYuanzhi Li, Ruosong Wang, Lin F. YangFOCS 2021 · 被引用 3 次
- Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount FactorJulien Grand-Clément, Marek PetrikNeurIPS 2023 · 被引用 25 次
- Horizon-free Learning for Markov Decision Processes and Games: Stochastically Bounded Rewards and Improved BoundsShengshi Li, Lin YangICML 2023 · 被引用 3 次
