A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time Bound
Gal Dalal, Balázs Szörényi, Gugan Thoppe
Abstract
Policy evaluation in reinforcement learning is often conducted using two-timescale stochastic approximation, which results in various gradient temporal difference methods such as GTD(0), GTD2, and TDC. Here, we provide convergence rate bounds for this suite of algorithms. Algorithms such as these have two iterates, and which are updated using two distinct stepsize sequences, and respectively. Assuming and with we show that, with high probability, the two iterates converge to their respective solutions and at rates given by and here, hides logarithmic terms. Via comparable lower bounds, we show that these bounds are, in fact, tight. To the best of our knowledge, ours is the first finite-time analysis which achieves these rates. While it was known that the two timescale components decouple asymptotically, our results depict this phenomenon more explicitly by showing that it in fact happens from some finite time onwards. Lastly, compared to existing works, our result applies to a broader family of stepsizes, including non-square summable ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 245df81a-6864-4787-9aa1-3b0a2906da06Cited by top-tier papers13
- CRPO: A New Approach for Safe Reinforcement Learning with Convergence GuaranteeTengyu Xu, Yingbin Liang, Guanghui LanICML 2021 · 171 citations
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General UtilitiesDonghao Ying, Yunkai Zhang, Yuhao Ding, Alec Koppel et al.NeurIPS 2023 · 28 citations
- A Single-timescale Analysis for Stochastic Approximation with Multiple Coupled SequencesHan Shen, Tianyi ChenNeurIPS 2022 · 25 citations
- Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function ApproximationGuojun Xiong, Jian LiNeurIPS 2023 · 23 citations
Related papers
- Gaussian Approximation for Two-Timescale Linear Stochastic ApproximationBogdan Butyrin, Artemy Rubtsov, Alexey Naumov, Vladimir V. Ulyanov et al.AAAI 2026 · 2 citations
- Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function ApproximationYue Wang, Shaofeng Zou, Yi ZhouNeurIPS 2021 · 12 citations
- Gradient Temporal Difference with Momentum: Stability and ConvergenceRohan Deb, Shalabh BhatnagarAAAI 2022 · 4 citations
- Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement LearningVagul Mahadevan, Claire Chen, Shuze D Liu, Shangtong ZhangICML 2026
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian SamplingHuaqing Xiong, Tengyu Xu, Yingbin Liang, Wei ZhangAAAI 2021 · 37 citations
