When to Sense and Control? A Time-adaptive Approach for Continuous-Time RL
Lenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler, Andreas Krause
摘要
Reinforcement learning (RL) excels in optimizing policies for discrete-time Markov decision processes (MDP). However, various systems are inherently continuous in time, making discrete-time MDPs an inexact modeling choice. In many applications, such as greenhouse control or medical treatments, each interaction (measurement or switching of action) involves manual intervention and thus is inherently costly. Therefore, we generally prefer a time-adaptive approach with fewer interactions with the system. In this work, we formalize an RL framework, Time-adaptive Control&Sensing (TaCoS), that tackles this challenge by optimizing over policies that besides control predict the duration of its application. Our formulation results in an extended MDP that any standard RL algorithm can solve. We demonstrate that state-of-the-art RL algorithms trained on TaCoS drastically reduce the interaction amount over their discrete-time counterpart while retaining the same or improved performance, and exhibiting robustness over discretization frequency. Finally, we propose OTaCoS, an efficient model-based algorithm for our setting. We show that OTaCoS enjoys sublinear regret for systems with sufficiently smooth dynamics and empirically results in further sample-efficiency gains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sample-efficient and Scalable Exploration in Continuous-Time RLKlemens Iten, Lenart Treven, Bhavya Sukhija, Florian Dörfler 等ICLR 2026 · 被引用 3 次
- Instance-Dependent Continuous-Time Reinforcement Learning via Maximum Likelihood EstimationRunze Zhao, Yue Yu, Ruhan Wang, Chunfeng Huang 等ICML 2026 · 被引用 1 次
- The Value of Sensory Information to a RobotArjun Krishna, Edward S. Hu, Dinesh JayaramanICLR 2025
它引用的顶会 Paper16
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi 等NeurIPS 2020 · 被引用 137 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 被引用 72 次
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes 等NeurIPS 2023 · 被引用 42 次
相关 Paper
- Efficient Exploration in Continuous-time Model-based Reinforcement LearningLenart Treven, Jonas Hübotter, Bhavya Sukhija, Florian Dörfler 等NeurIPS 2023 · 被引用 24 次
- Model-based Reinforcement Learning for Semi-Markov Decision Processes with Neural ODEsJianzhun Du, Joseph Futoma, Finale Doshi-VelezNeurIPS 2020 · 被引用 63 次
- Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-offZichen Vincent Zhang, Johannes Kirschner, Junxi Zhang, Francesco Zanini 等NeurIPS 2023 · 被引用 3 次
- Logarithmic Regret Bound in Partially Observable Linear Dynamical SystemsSahin Lale, Kamyar Azizzadenesheli, Babak Hassibi, Anima AnandkumarNeurIPS 2020 · 被引用 106 次
- Improving planning and MBRL with temporally-extended actionsPalash Chatterjee, Roni KhardonNeurIPS 2025
