An Experimental Design Perspective on Model-Based Reinforcement Learning
Viraj Mehta, Biswajit Paria, Jeff Schneider, Stefano Ermon, Willie Neiswanger
摘要
In many practical applications of RL, it is expensive to observe state transitions from the environment. For example, in the problem of plasma control for nuclear fusion, computing the next state for a given state-action pair requires querying an expensive transition function which can lead to many hours of computer simulation or dollars of scientific research. Such expensive data collection prohibits application of standard RL algorithms which usually require a large number of observations to learn. In this work, we address the problem of efficiently learning a policy while making a minimal number of state-action queries to the transition function. In particular, we leverage ideas from Bayesian optimal experimental design to guide the selection of state-action queries for efficient learning. We propose an acquisition function that quantifies how much information a state-action pair would provide about the optimal solution to a Markov decision process. At each iteration, our algorithm maximizes this acquisition function, to choose the most informative state-action pair to be queried, thus yielding a data-efficient RL approach. We experiment with a variety of simulated continuous control problems and show that our approach learns an optimal policy with up to -- less data than model-based RL baselines and -- less data than model-free RL baselines. We also provide several ablated comparisons which point to substantial improvements arising from the principled method of obtaining data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes 等NeurIPS 2023 · 被引用 42 次
- Exploration via Planning for Information about the Optimal TrajectoryViraj Mehta, Ian Char, Joseph Abbate, Rory Conlin 等NeurIPS 2022 · 被引用 12 次
- Episodic Future Thinking Mechanism for Multi-agent Reinforcement LearningDongsu Lee, Minhae KwonNeurIPS 2024 · 被引用 9 次
- PID-Inspired Inductive Biases for Deep Reinforcement Learning in Partially Observable Control TasksIan Char, Jeff SchneiderNeurIPS 2023 · 被引用 8 次
- DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under UncertaintyMingxuan Cui, Duo Zhou, Yuxuan Han, Grani A. Hanasusanto 等ICLR 2026 · 被引用 6 次
它引用的顶会 Paper7
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky 等ICML 2020 · 被引用 186 次
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 被引用 164 次
- Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative ModelGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu 等NeurIPS 2020 · 被引用 159 次
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski 等ICML 2020 · 被引用 52 次
- Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual InformationWillie Neiswanger, Ke Alexander Wang, Stefano ErmonICML 2021 · 被引用 40 次
相关 Paper
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr 等ICLR 2020 · 被引用 104 次
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek 等NeurIPS 2021 · 被引用 27 次
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 被引用 67 次
- ALINE: Joint Amortization for Bayesian Inference and Active Data AcquisitionDaolang Huang, Xinyi Wen, Ayush Bharti, Samuel Kaski 等NeurIPS 2025 · 被引用 8 次
- Transition Constrained Bayesian Optimization via Markov Decision ProcessesJose Pablo Folch, Calvin Tsay, Robert M. Lee, Behrang Shafei 等NeurIPS 2024 · 被引用 10 次
