Policy Evaluation for Variance in Average Reward Reinforcement Learning
Shubhada Agrawal, Prashanth L. A., Siva Theja Maguluri
摘要
We consider an average reward reinforcement learning (RL) problem and work with asymptotic variance as a risk measure to model safety-critical applications. We design a temporal-difference (TD) type algorithm tailored for policy evaluation in this context. Our algorithm is based on linear stochastic approximation of an equivalent formulation of the asymptotic variance in terms of the solution of the Poisson equation. We consider both the tabular and linear function approximation settings, and establish Õ(1/k) finite time convergence rate, where k is the number of steps of the algorithm. Our work paves the way for developing actor-critic style algorithms for varianceconstrained RL. To the best of our knowledge, our result provides the first sequential estimator for asymptotic variance of a Markov chain with provable finite sample guarantees, which is of independent interest.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Least Squares Regression with Markovian Data: Fundamental Limits and AlgorithmsDheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain 等NeurIPS 2020 · 被引用 73 次
- Finite Sample Analysis of Average-Reward TD Learning and -LearningSheng Zhang, Zhe Zhang, Siva Theja MaguluriNeurIPS 2021 · 被引用 48 次
- Bootstrapping Fitted Q-Evaluation for Off-Policy InferenceBotao Hao, Xiang Ji, Yaqi Duan, Hao Lu 等ICML 2021 · 被引用 46 次
相关 Paper
- Variance Penalized On-Policy and Off-Policy Actor-CriticArushi Jain, Gandharv Patil, Ayush Jain, Khimya Khetarpal 等AAAI 2021 · 被引用 11 次
- Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian NoiseEthan Blaser, Shangtong ZhangAAAI 2026 · 被引用 3 次
- Reanalysis of Variance Reduced Temporal Difference LearningTengyu Xu, Zhe Wang, Yi Zhou, Yingbin LiangICLR 2020 · 被引用 46 次
- Efficient Policy Evaluation with Safety Constraint for Reinforcement LearningClaire Chen, Shuze Daniel Liu, Shangtong ZhangICLR 2025
- Finite Sample Analysis of Linear Temporal Difference Learning with Arbitrary FeaturesZixuan Xie, Xinyu Liu, Rohan Chandra, Shangtong ZhangNeurIPS 2025 · 被引用 6 次
