Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-off
Zichen Vincent Zhang, Johannes Kirschner, Junxi Zhang, Francesco Zanini, Alex Ayoub, Masood Dehghan, Dale Schuurmans
Abstract
A default assumption in reinforcement learning (RL) and optimal control is that observations arrive at discrete time points on a fixed clock cycle. Yet, many applications involve continuous-time systems where the time discretization, in principle, can be managed. The impact of time discretization on RL methods has not been fully characterized in existing theory, but a more detailed analysis of its effect could reveal opportunities for improving data-efficiency. We address this gap by analyzing Monte-Carlo policy evaluation for LQR systems and uncover a fundamental trade-off between approximation and statistical error in value estimation. Importantly, these two errors behave differently to time discretization, leading to an optimal choice of temporal resolution for a given data budget. These findings show that managing the temporal resolution can provably improve policy evaluation efficiency in LQR systems with finite data. Empirically, we demonstrate the trade-off in numerical simulations of LQR instances and standard RL benchmarks for non-linear continuous control.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Continuous-time Model-based Reinforcement LearningÇagatay Yildiz, Markus Heinonen, Harri LähdesmäkiICML 2021 · 72 citations
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni et al.ICML 2020 · 43 citations
- Value Iteration in Continuous Actions, States and TimeMichael Lutter, Shie Mannor, Jan Peters, Dieter Fox et al.ICML 2021 · 40 citations
- Time Discretization-Invariant Safe Action Repetition for Policy Gradient MethodsSeohong Park, Jaekyeom Kim, Gunhee KimNeurIPS 2021 · 33 citations
- Provably adaptive reinforcement learning in metric spacesTongyi Cao, Akshay KrishnamurthyNeurIPS 2020 · 8 citations
Related papers
- Robust On-Policy Sampling for Data-Efficient Policy Evaluation in Reinforcement LearningRujie Zhong, Duohan Zhang, Lukas Schäfer, Stefano V. Albrecht et al.NeurIPS 2022 · 19 citations
- The surprising efficiency of temporal difference learning for rare event predictionXiaoou Cheng, Jonathan WeareNeurIPS 2024 · 10 citations
- When to Sense and Control? A Time-adaptive Approach for Continuous-Time RLLenart Treven, Bhavya Sukhija, Yarden As, Florian Dörfler et al.NeurIPS 2024 · 10 citations
- Efficient RL with Impaired Observability: Learning to Act with Delayed and Missing State ObservationsMinshuo Chen, Yu Bai, H. Vincent Poor, Mengdi WangNeurIPS 2023 · 19 citations
- Improving planning and MBRL with temporally-extended actionsPalash Chatterjee, Roni KhardonNeurIPS 2025
