Integral Performance Approximation for Continuous-Time Reinforcement Learning Control
Brent A. Wallace, Jennie Si
摘要
We introduce integral performance approximation (IPA), a new continuous-time reinforcement learning (CT-RL) control method. It leverages an affine nonlinear dynamic model, which partially captures the dynamics of the physical environment, alongside state-action trajectory data to enable optimal control with great data efficiency and robust control performance. Utilizing Kleinman algorithm structures allows IPA to provide theoretical guarantees of learning convergence, solution optimality, and closed-loop stability. Furthermore, we demonstrate the effectiveness of IPA on three CT-RL environments including hypersonic vehicle (HSV) control, which has additional challenges caused by unstable and nonminimum phase dynamics. As a result, we demonstrate that the IPA method leads to new, SOTA control design and performance in CT-RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 被引用 72 次
- Value Iteration in Continuous Actions, States and TimeMichael Lutter, Shie Mannor, Jan Peters, Dieter Fox 等ICML 2021 · 被引用 40 次
- Impact of Computation in Integral Reinforcement Learning for Continuous-Time ControlWenhan Cao, Wei PanICLR 2024 · 被引用 1 次
相关 Paper
- Model-based Reinforcement Learning for Parameterized Action SpacesRenhao Zhang, Haotian Fu, Yilin Miao, George KonidarisICML 2024 · 被引用 8 次
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 2 次
- Continuous-Time Value Iteration for Multi-Agent Reinforcement LearningXuefeng Wang, Lei Zhang, Henglin Pu, Ahmed Hussain Qureshi 等ICLR 2026 · 被引用 3 次
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi 等NeurIPS 2020 · 被引用 137 次
- Safe Continuous-time Multi-Agent Reinforcement Learning via Epigraph FormXuefeng Wang, Lei Zhang, Henglin Pu, Husheng Li 等ICLR 2026 · 被引用 1 次
