Toward Efficient Gradient-Based Value Estimation
Arsalan Sharifnassab, Richard S. Sutton
Abstract
Gradient-based methods for value estimation in reinforcement learning have favorable stability properties, but they are typically much slower than Temporal Difference (TD) learning methods. We study the root causes of this slowness and show that Mean Square Bellman Error (MSBE) is an ill-conditioned loss function in the sense that its Hessian has large condition-number. To resolve the adverse effect of poor conditioning of MSBE on gradient based methods, we propose a low complexity batch-free proximal method that approximately follows the Gauss-Newton direction and is asymptotically robust to parameterization. Our main algorithm, called RANS, is efficient in the sense that it is significantly faster than the residual gradient methods while having almost the same computational complexity, and is competitive with TD on the classic problems that we tested.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f57f950-14d0-4a9b-b1cb-6dba9e80ae12Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta et al.ICML 2020 · 49 citations
- Geometric Insights into the Convergence of Nonlinear TD LearningDavid Brandfonbrener, Joan BrunaICLR 2020 · 18 citations
- Gradient Temporal Difference with Momentum: Stability and ConvergenceRohan Deb, Shalabh BhatnagarAAAI 2022 · 4 citations
Related papers
- Statistically Efficient Off-Policy Policy GradientsNathan Kallus, Masatoshi UeharaICML 2020 · 43 citations
- Generalizing Gaussian Smoothing for Random SearchKatelyn Gao, Ozan SenerICML 2022 · 22 citations
- Reducing Sampling Error in Batch Temporal Difference LearningBrahma S. Pavse, Ishan Durugkar, Josiah Hanna, Peter StoneICML 2020 · 14 citations
- Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement LearningMotoki Omura, Takayuki Osa, Yusuke Mukuta, Tatsuya HaradaAAAI 2024 · 1 citation
- Learning Near Optimal Policies with Low Inherent Bellman ErrorAndrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma BrunskillICML 2020 · 238 citations
