Finite-Sample Analysis of Contractive Stochastic Approximation Using Smooth Convex Envelopes
Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan Shanmugam
摘要
Stochastic Approximation (SA) is a popular approach for solving fixed-point equations where the information is corrupted by noise. In this paper, we consider an SA involving a contraction mapping with respect to an arbitrary norm, and show its finite-sample error bounds while using different stepsizes. The idea is to construct a smooth Lyapunov function using the generalized Moreau envelope, and show that the iterates of SA have negative drift with respect to that Lyapunov function. Our result is applicable in Reinforcement Learning (RL). In particular, we use it to establish the first-known convergence rate of the V-trace algorithm for off-policy TD-learning [18] . Importantly, our construction results in only a logarithmic dependence of the convergence bound on the size of the state-space.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- The Blessing of Heterogeneity in Federated Q-Learning: Linear Speedup and BeyondJiin Woo, Gauri Joshi, Yuejie ChiICML 2023 · 被引用 36 次
- Finite-Sample Analysis of Off-Policy Natural Actor-Critic AlgorithmSajad Khodadadian, Zaiwei Chen, Siva Theja MaguluriICML 2021 · 被引用 33 次
- Online Target Q-learning with Reverse Experience Replay: Efficiently finding the Optimal Policy for Linear MDPsNaman Agarwal, Syomantak Chaudhuri, Prateek Jain, Dheeraj Mysore Nagaraj 等ICLR 2022 · 被引用 24 次
- A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic GamesZaiwei Chen, Kaiqing Zhang, Eric Mazumdar, Asuman E. Ozdaglar 等NeurIPS 2023 · 被引用 22 次
- Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman OperatorsZaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan ShanmugamNeurIPS 2021 · 被引用 21 次
它引用的顶会 Paper1
相关 Paper
- Asymptotic and Finite Sample Analysis of Nonexpansive Stochastic Approximations with Markovian NoiseEthan Blaser, Shangtong ZhangAAAI 2026 · 被引用 3 次
- Finite Sample Analysis of Average-Reward TD Learning and -LearningSheng Zhang, Zhe Zhang, Siva Theja MaguluriNeurIPS 2021 · 被引用 48 次
- Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement LearningVagul Mahadevan, Claire Chen, Shuze D Liu, Shangtong ZhangICML 2026
- On Convergence of Gradient Expected Sarsa(λ)Long Yang, Gang Zheng, Yu Zhang, Qian Zheng 等AAAI 2021 · 被引用 5 次
- A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak AveragingSajad Khodadadian, Martin ZubeldiaNeurIPS 2025 · 被引用 4 次
