Exploration via Planning for Information about the Optimal Trajectory
Viraj Mehta, Ian Char, Joseph Abbate, Rory Conlin, Mark D. Boyer, Stefano Ermon, Jeff Schneider, Willie Neiswanger
Abstract
Many potential applications of reinforcement learning (RL) are stymied by the large numbers of samples required to learn an effective policy. This is especially true when applying RL to real-world control tasks, e.g. in the sciences or robotics, where executing a policy in the environment is costly. In popular RL algorithms, agents typically explore either by adding stochasticity to a reward-maximizing policy or by attempting to gather maximal information about environment dynamics without taking the given task into account. In this work, we develop a method that allows us to plan for exploration while taking both the task and the current knowledge about the dynamics into account. The key insight to our approach is to plan an action sequence that maximizes the expected information gain about the optimal trajectory for the task at hand. We demonstrate that our method learns strong policies with 2x fewer samples than strong exploration baselines and 200x fewer samples than model free methods on a diverse set of low-to-medium dimensional control tasks in both the open-loop and closed-loop control settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- PID-Inspired Inductive Biases for Deep Reinforcement Learning in Partially Observable Control TasksIan Char, Jeff SchneiderNeurIPS 2023 · 8 citations
- Sample-efficient Bayesian Optimisation Using Known InvariancesTheodore Brown, Alexandru Cioba, Ilija BogunovicNeurIPS 2024 · 6 citations
- Sampling-based Multi-dimensional RecalibrationYoungseog Chung, Ian Char, Jeff SchneiderICML 2024 · 4 citations
- Near-optimal Policy Identification in Active Reinforcement LearningXiang Li, Viraj Mehta, Johannes Kirschner, Ian Char et al.ICLR 2023
- Efficient Model-Based Reinforcement Learning Through Optimistic Thompson SamplingJasmine Bayrooti, Carl Henrik Ek, Amanda ProrokICLR 2025
Builds on9
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky et al.ICML 2020 · 186 citations
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 120 citations
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski et al.ICML 2020 · 52 citations
Related papers
- FLEX: an Adaptive Exploration Algorithm for Nonlinear SystemsMatthieu Blanke, Marc LelargeICML 2023 · 5 citations
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 56 citations
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel et al.ICLR 2025
- Task-Optimal Exploration in Linear Dynamical SystemsAndrew J. Wagenmaker, Max Simchowitz, Kevin JamiesonICML 2021 · 24 citations
- Reinforcement Learning with Simple Sequence PriorsTankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz et al.NeurIPS 2023 · 18 citations
