Impact of Computation in Integral Reinforcement Learning for Continuous-Time Control
Wenhan Cao, Wei Pan
Abstract
Integral reinforcement learning (IntRL) demands the precise computation of the utility function's integral at its policy evaluation (PEV) stage. This is achieved through quadrature rules, which are weighted sums of utility functions evaluated from state samples obtained in discrete time. Our research reveals a critical yet underexplored phenomenon: the choice of the computational method -- in this case, the quadrature rule -- can significantly impact control performance. This impact is traced back to the fact that computational errors introduced in the PEV stage can affect the policy iteration's convergence behavior, which in turn affects the learned controller. To elucidate how computation impacts control, we draw a parallel between IntRL's policy iteration and Newton's method applied to the Hamilton-Jacobi-Bellman equation. In this light, computational error in PEV manifests as an extra error term in each iteration of Newton's method, with its upper bound proportional to the computational error. Further, we demonstrate that when the utility function resides in a reproducing kernel Hilbert space (RKHS), the optimal quadrature is achievable by employing Bayesian quadrature with the RKHS-inducing kernel function. We prove that the local convergence rates for IntRL using the trapezoidal rule and Bayesian quadrature with a Matérn kernel to be and , where is the number of evenly-spaced samples and is the Matérn kernel's smoothness parameter. These theoretical findings are finally validated by two canonical control tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64f2ee4c-4186-4f26-9dfc-798da738128aCited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Policy Newton Algorithm in Reproducing Kernel Hilbert SpaceYixian Zhang, Huaze Tang, Changxu Wei, Chao Wang et al.ICLR 2026 · 3 citations
- A Non-asymptotic Analysis of Non-parametric Temporal-Difference LearningEloïse Berthier, Ziad Kobeissi, Francis R. BachNeurIPS 2022 · 6 citations
- Kernel Quadrature with Randomly Pivoted CholeskyEthan Epperly, Elvira MorenoNeurIPS 2023 · 16 citations
- Kernelized Reinforcement Learning with Order Optimal Regret BoundsSattar Vakili, Julia OlkhovskayaNeurIPS 2023 · 22 citations
- Error Propagation in Dynamic Programming: From Stochastic Control to American Option PricingAndrea Della Vecchia, Damir FilipovicICML 2026
