Lune

AAAI2021顶会

Scalable First-Order Methods for Robust MDPs

Julien Grand-Clément, Christian Kroer

2021年份
33被引次数
15顶会引用

摘要

Robust Markov Decision Processes (MDPs) are a powerful framework for modeling sequential decision making problems with model uncertainty. This paper proposes the first first-order framework for solving robust MDPs. Our algorithm interleaves primal-dual first-order updates with approximate Value Iteration updates. By carefully controlling the tradeoff between the accuracy and cost of Value Iteration updates, we achieve an ergodic convergence rate of O A 2 S 3 log(S) log( -1 ) -1 for the best choice of parameters on ellipsoidal and Kullback-Leibler s-rectangular uncertainty sets, where S and A is the number of states and actions, respectively. Our dependence on the number of states and actions is significantly better (by a factor of O(A 1.5 S 1.5 )) than that of pure Value Iteration algorithms. In numerical experiments on ellipsoidal uncertainty sets we show that our algorithm is significantly more scalable than state-of-the-art approaches. Our framework is also the first one to solve robust MDPs with s-rectangular KL uncertainty sets.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper15

问问它们各自怎么用它

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖