Integral Performance Approximation for Continuous-Time Reinforcement Learning Control
Brent A. Wallace, Jennie Si
Abstract
We introduce integral performance approximation (IPA), a new continuous-time reinforcement learning (CT-RL) control method. It leverages an affine nonlinear dynamic model, which partially captures the dynamics of the physical environment, alongside state-action trajectory data to enable optimal control with great data efficiency and robust control performance. Utilizing Kleinman algorithm structures allows IPA to provide theoretical guarantees of learning convergence, solution optimality, and closed-loop stability. Furthermore, we demonstrate the effectiveness of IPA on three CT-RL environments including hypersonic vehicle (HSV) control, which has additional challenges caused by unstable and nonminimum phase dynamics. As a result, we demonstrate that the IPA method leads to new, SOTA control design and performance in CT-RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c396882-2729-481d-9c0a-93ec696daa62Builds on3
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 72 citations
- Value Iteration in Continuous Actions, States and TimeMichael Lutter, Shie Mannor, Jan Peters, Dieter Fox et al.ICML 2021 · 40 citations
- Impact of Computation in Integral Reinforcement Learning for Continuous-Time ControlWenhan Cao, Wei PanICLR 2024 · 1 citation
Related papers
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 8 citations
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 2 citations
- Continuous-Time Value Iteration for Multi-Agent Reinforcement LearningXuefeng Wang, Lei Zhang, Henglin Pu, Ahmed Hussain Qureshi et al.ICLR 2026 · 3 citations
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi et al.NeurIPS 2020 · 137 citations
- Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph FormXuefeng Wang, Lei Zhang, Henglin Pu, Husheng Li et al.ICLR 2026 · 1 citation
