Lune

NeurIPS2022顶会

Provably Efficient Offline Multi-agent Reinforcement Learning via Strategy-wise Bonus

Qiwen Cui, Simon S. Du

2022年份
34被引次数
16顶会引用

摘要

This paper considers offline multi-agent reinforcement learning. We propose the strategy-wise concentration principle which directly builds a confidence interval for the joint strategy, in contrast to the point-wise concentration principle that builds a confidence interval for each point in the joint action space. For two-player zero-sum Markov games, by exploiting the convexity of the strategy-wise bonus, we propose a computationally efficient algorithm whose sample complexity enjoys a better dependency on the number of actions than the prior methods based on the point-wise bonus. Furthermore, for offline multi-agent general-sum Markov games, based on the strategy-wise bonus and a novel surrogate function, we give the first algorithm whose sample complexity only scales ∑i=1mAi\sum_{i=1}^mA_i where AiA_i is the action size of the ii-th player and mm is the number of players. In sharp contrast, the sample complexity of methods based on the point-wise bonus would scale with the size of the joint action space Πi=1mAi\Pi_{i=1}^m A_i due to the curse of multiagents. Lastly, all of our algorithms can naturally take a pre-specified strategy class Π\Pi as input and output a strategy that is close to the best strategy in Π\Pi. In this setting, the sample complexity only scales with log⁡∣Π∣\log |\Pi| instead of ∑i=1mAi\sum_{i=1}^mA_i.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper16

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖