Exploration via Planning for Information about the Optimal Trajectory
Viraj Mehta, Ian Char, Joseph Abbate, Rory Conlin, Mark D. Boyer, Stefano Ermon, Jeff Schneider, Willie Neiswanger
摘要
Many potential applications of reinforcement learning (RL) are stymied by the large numbers of samples required to learn an effective policy. This is especially true when applying RL to real-world control tasks, e.g. in the sciences or robotics, where executing a policy in the environment is costly. In popular RL algorithms, agents typically explore either by adding stochasticity to a reward-maximizing policy or by attempting to gather maximal information about environment dynamics without taking the given task into account. In this work, we develop a method that allows us to plan for exploration while taking both the task and the current knowledge about the dynamics into account. The key insight to our approach is to plan an action sequence that maximizes the expected information gain about the optimal trajectory for the task at hand. We demonstrate that our method learns strong policies with 2x fewer samples than strong exploration baselines and 200x fewer samples than model free methods on a diverse set of low-to-medium dimensional control tasks in both the open-loop and closed-loop control settings. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PID-Inspired Inductive Biases for Deep Reinforcement Learning in Partially Observable Control TasksIan Char, Jeff SchneiderNeurIPS 2023 · 被引用 8 次
- Sample-efficient Bayesian Optimisation Using Known InvariancesTheodore Brown, Alexandru Cioba, Ilija BogunovicNeurIPS 2024 · 被引用 6 次
- Sampling-based Multi-dimensional RecalibrationYoungseog Chung, Ian Char, Jeff SchneiderICML 2024 · 被引用 4 次
- Near-optimal Policy Identification in Active Reinforcement LearningXiang Li, Viraj Mehta, Johannes Kirschner, Ian Char 等ICLR 2023
- Efficient Model-Based Reinforcement Learning Through Optimistic Thompson SamplingJasmine Bayrooti, Carl Henrik Ek, Amanda ProrokICLR 2025
它引用的顶会 Paper9
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 被引用 239 次
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky 等ICML 2020 · 被引用 186 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski 等ICML 2020 · 被引用 52 次
相关 Paper
- FLEX: an Adaptive Exploration Algorithm for Nonlinear SystemsMatthieu Blanke, Marc LelargeICML 2023 · 被引用 5 次
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 被引用 56 次
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel 等ICLR 2025
- Task-Optimal Exploration in Linear Dynamical SystemsAndrew J. Wagenmaker, Max Simchowitz, Kevin JamiesonICML 2021 · 被引用 24 次
- Reinforcement Learning with Simple Sequence PriorsTankred Saanum, Noémi Élteto, Peter Dayan, Marcel Binz 等NeurIPS 2023 · 被引用 18 次
