NeoRL: Efficient Exploration for Nonepisodic RL
Bhavya Sukhija, Lenart Treven, Florian Dörfler, Stelian Coros, Andreas Krause
Abstract
We study the problem of nonepisodic reinforcement learning (RL) for nonlinear dynamical systems, where the system dynamics are unknown and the RL agent has to learn from a single trajectory, i.e., without resets. We propose Nonepisodic Optimistic RL (NeoRL), an approach based on the principle of optimism in the face of uncertainty. NeoRL uses well-calibrated probabilistic models and plans optimistically w.r.t. the epistemic uncertainty about the unknown dynamics. Under continuity and bounded energy assumptions on the system, we provide a first-of-its-kind regret bound of for general nonlinear systems with Gaussian process dynamics. We compare NeoRL to other baselines on several deep RL environments and empirically demonstrate that NeoRL achieves the optimal average cost while incurring the least regret.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76aba389-5bb9-4e14-b26d-e285e8ccecacCited by top-tier papers4
- SOMBRL: Scalable and Optimistic Model-Based RLBhavya Sukhija, Lenart Treven, Carmelo Sferrazza, Florian Dörfler et al.NeurIPS 2025 · 9 citations
- No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian ProcessesJasmine Bayrooti, Sattar Vakili, Amanda Prorok, Carl Henrik EkNeurIPS 2025 · 5 citations
- Policy Search via Bayesian Optimization with Temporal Difference Gaussian ProcessesArmin Lederer, Anuj Srivastava, Marco Bagatella, Andreas KrauseICML 2026
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel et al.ICLR 2025
Builds on19
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Naive Exploration is Optimal for Online LQRMax Simchowitz, Dylan J. FosterICML 2020 · 209 citations
- The Ingredients of Real World Robotic Reinforcement LearningHenry Zhu, Justin Yu, Abhishek Gupta, Dhruv Shah et al.ICLR 2020 · 202 citations
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi et al.NeurIPS 2020 · 137 citations
Related papers
- Efficient Exploration in Continuous-time Model-based Reinforcement LearningLenart Treven, Jonas Hübotter, Bhavya Sukhija, Florian Dörfler et al.NeurIPS 2023 · 24 citations
- Sample-efficient and Scalable Exploration in Continuous-Time RLKlemens Iten, Lenart Treven, Bhavya Sukhija, Florian Dörfler et al.ICLR 2026 · 3 citations
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes et al.NeurIPS 2023 · 42 citations
- Dynamic Regret of Adversarial MDPs with Unknown Transition and Linear Function ApproximationLong-Fei Li, Peng Zhao, Zhi-Hua ZhouAAAI 2024 · 3 citations
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar et al.NeurIPS 2021 · 110 citations
