Lune

ICML2023顶会

Horizon-free Learning for Markov Decision Processes and Games: Stochastically Bounded Rewards and Improved Bounds

Shengshi Li, Lin Yang

出版方
2023年份
3被引次数
1顶会引用

摘要

Horizon dependence is an important difference between reinforcement learning and other machine learning paradigms. Yet, existing results tackling the (exact) horizon dependence either assume that the reward is bounded per step, introducing unfair comparison, or assume strict total boundedness that requires the sum of rewards to be bounded almost surely -allowing only restricted noise on the reward observation. This paper addresses these limitations by introducing a new relaxationexpected boundedness on rewards, where we allow the reward to be stochastic with only boundedness on the expected sum -opening the door to study horizon-dependence with a much broader set of reward functions with noises. We establish a novel generic algorithm that achieves nohorizon dependence in terms of sample complexity for both Markov Decision Processes (MDP) and Games, via reduction to a good-conditioned auxiliary Markovian environment, in which only "important" state-action pairs are preserved. The algorithm takes only Õ( S 2 A ✏ 2 ) episodes interacting with such an environment to achieve an ✏-optimal policy/strategy (with high probability), improving (Zhang et al., 2022) (which only applies to MDPs with deterministic rewards). Here S is the number of states and A is the number of actions, and the bound is independent of the horizon H.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖