A Temporal Difference Method for Stochastic Continuous Dynamics
Haruki Settai, Naoya Takeishi, Takehisa Yairi
摘要
For continuous systems modeled by dynamical equations such as ODEs and SDEs, Bellman's Principle of Optimality takes the form of the Hamilton-Jacobi-Bellman (HJB) equation, which provides the theoretical target of reinforcement learning (RL). Although recent advances in RL successfully leverage this formulation, the existing methods typically assume the underlying dynamics are known a priori because they need explicit access to the coefficient functions of dynamical equations to update the value function following the HJB equation. We address this inherent limitation of HJB-based RL; we propose a model-free approach still targeting the HJB equation and propose the corresponding temporal difference method. We establish exponential convergence of the idealized continuous-time dynamics and empirically demonstrate its potential advantages over transition-kernel-based formulations. The proposed formulation paves the way toward bridging stochastic control and model-free reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 被引用 96 次
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 被引用 72 次
- Policy Optimization for Continuous Reinforcement LearningHanyang Zhao, Wenpin Tang, David D. YaoNeurIPS 2023 · 被引用 47 次
相关 Paper
- Sample-efficient and Scalable Exploration in Continuous-Time RLKlemens Iten, Lenart Treven, Bhavya Sukhija, Florian Dörfler 等ICLR 2026 · 被引用 3 次
- Enhancing Value Function Estimation through First-Order State-Action Dynamics in Offline Reinforcement LearningYun-Hsuan Lien, Ping-Chun Hsieh, Tzu-Mao Li, Yu-Shuen WangICML 2024 · 被引用 4 次
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEsJianzhun Du, Joseph Futoma, Finale Doshi-VelezNeurIPS 2020 · 被引用 63 次
- Efficient Exploration in Continuous-time Model-based Reinforcement LearningLenart Treven, Jonas Hübotter, Bhavya Sukhija, Florian Dörfler 等NeurIPS 2023 · 被引用 24 次
- Reinforcement Learning with Non-Exponential DiscountingMatthias Schultheis, Constantin A. Rothkopf, Heinz KoepplNeurIPS 2022 · 被引用 19 次
