Lune

NeurIPS2025顶会

Regret Analysis of Average-Reward Unichain MDPs via an Actor-Critic Approach

Swetha Ganesh, Vaneet Aggarwal

2025年份
9被引次数
2顶会引用

摘要

Actor-Critic methods are widely used for their scalability, yet existing theoretical guarantees for infinite-horizon average-reward Markov Decision Processes (MDPs) often rely on restrictive ergodicity assumptions. We propose NAC-B, a Natural Actor-Critic with Batching, that achieves order-optimal regret of Õ( √ T ) in infinitehorizon average-reward MDPs under the unichain assumption, which permits both transient states and periodicity. This assumption is among the weakest under which the classic policy gradient theorem remains valid for average-reward settings. NAC-B employs function approximation for both the actor and the critic, enabling scalability to problems with large state and action spaces. The use of batching in our algorithm helps mitigate potential periodicity in the MDP and reduces stochasticity in gradient estimates, and our analysis formalizes these benefits through the introduction of the constants C hit and C tar , which characterize the rate at which empirical averages over Markovian samples converge to the stationary distribution. 39th Conference on Neural Information Processing Systems (NeurIPS 2025). Algorithm Regret Ergodicity-free General Policy MDP-OOMD [38] Õ( √ T ) No No Optimistic Q-learning [38] Õ(T 2/3 ) Yes (1) No MDP-EXP2 [39] Õ( √ T ) No No UCB-AVG [45] Õ( √ T ) Yes (1) No PPG [9] Õ(T 3/4 ) No Yes PHAPG [17] Õ( √ T ) No Yes Optimistic Q-learning [3] Õ( √ T ) Yes No γ-DC-LSCVI-UCB [20] Õ( √ T ) Yes (1) No This work (Algorithm 1) Õ( √ T ) Yes Yes

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper16

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖