Lune

ICLR2023Top-tier venue

On the Performance of Temporal Difference Learning With Neural Networks

Haoxing Tian, Ioannis Ch. Paschalidis, Alex Olshevsky

2023Year
2Top-tier citations

Abstract

Neural Temporal Difference (TD) Learning is an approximate temporal difference method for policy evaluation that uses a neural network for function approximation. Analysis of Neural TD Learning has proven to be challenging. In this paper we provide a convergence analysis of Neural TD Learning with a projection onto B(θ 0 , ω), a ball of fixed radius ω around the initial point θ 0 . We show an approximation bound of O(ϵ) + Õ(1/ √ m) where ϵ is the approximation quality of the best neural network in B(θ 0 , ω) and m is the width of all hidden layers in the network.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6235113d-2128-46bb-bc96-9cbe98ea0d95

Cited by top-tier papers2

Ask how each one uses it

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines