A Non-asymptotic Analysis of Non-parametric Temporal-Difference Learning
Eloïse Berthier, Ziad Kobeissi, Francis R. Bach
Abstract
Temporal-difference learning is a popular algorithm for policy evaluation. In this paper, we study the convergence of the regularized non-parametric TD(0) algorithm, in both the independent and Markovian observation settings. In particular, when TD is performed in a universal reproducing kernel Hilbert space (RKHS), we prove convergence of the averaged iterates to the optimal value function, even when it does not belong to the RKHS. We provide explicit convergence rates that depend on a source condition relating the regularity of the optimal value function to the RKHS. We illustrate this convergence numerically on a simple continuous-state Markov reward process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da6ca0cd-5445-4f47-89bd-1ecf631a667fCited by top-tier papers2
- On Double Descent in Reinforcement Learning with LSTD and Random FeaturesDavid Brellmann, Eloïse Berthier, David Filliat, Goran FrehseICLR 2024 · 2 citations
- Loss Dynamics of Temporal Difference Reinforcement LearningBlake Bordelon, Paul Masset, Henry Kuo, Cengiz PehlevanNeurIPS 2023
Builds on4
- Least Squares Regression with Markovian Data: Fundamental Limits and AlgorithmsDheeraj Nagaraj, Xian Wu, Guy Bresler, Prateek Jain et al.NeurIPS 2020 · 73 citations
- Tight Nonparametric Convergence Rates for Stochastic Gradient Descent under the Noiseless Linear ModelRaphaël Berthier, Francis R. Bach, Pierre GaillardNeurIPS 2020 · 49 citations
- Reanalysis of Variance Reduced Temporal Difference LearningTengyu Xu, Zhe Wang, Yi Zhou, Yingbin LiangICLR 2020 · 46 citations
- Kernel-Based Reinforcement Learning: A Finite-Time AnalysisOmar Darwiche Domingues, Pierre Ménard, Matteo Pirotta, Emilie Kaufmann et al.ICML 2021 · 24 citations
Related papers
- Sampling Complexity of TD and PPO in RKHSLU ZOU, Wendi Ren, WEIZHONG ZHANG, Liang Ding et al.ICLR 2026 · 1 citation
- Policy Newton Algorithm in Reproducing Kernel Hilbert SpaceYixian Zhang, Huaze Tang, Changxu Wei, Chao Wang et al.ICLR 2026 · 3 citations
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 2 citations
- On the Performance of Temporal Difference Learning With Neural NetworksHaoxing Tian, Ioannis Ch. Paschalidis, Alex OlshevskyICLR 2023
- Kernelized Reinforcement Learning with Order Optimal Regret BoundsSattar Vakili, Julia OlkhovskayaNeurIPS 2023 · 22 citations
