An Experimental Design Perspective on Model-Based Reinforcement Learning
Viraj Mehta, Biswajit Paria, Jeff Schneider, Stefano Ermon, Willie Neiswanger
Abstract
In many practical applications of RL, it is expensive to observe state transitions from the environment. For example, in the problem of plasma control for nuclear fusion, computing the next state for a given state-action pair requires querying an expensive transition function which can lead to many hours of computer simulation or dollars of scientific research. Such expensive data collection prohibits application of standard RL algorithms which usually require a large number of observations to learn. In this work, we address the problem of efficiently learning a policy while making a minimal number of state-action queries to the transition function. In particular, we leverage ideas from Bayesian optimal experimental design to guide the selection of state-action queries for efficient learning. We propose an acquisition function that quantifies how much information a state-action pair would provide about the optimal solution to a Markov decision process. At each iteration, our algorithm maximizes this acquisition function, to choose the most informative state-action pair to be queried, thus yielding a data-efficient RL approach. We experiment with a variety of simulated continuous control problems and show that our approach learns an optimal policy with up to -- less data than model-based RL baselines and -- less data than model-free RL baselines. We also provide several ablated comparisons which point to substantial improvements arising from the principled method of obtaining data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82240352-ca0a-4b30-a3a3-2b09aab52720Cited by top-tier papers9
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes et al.NeurIPS 2023 · 42 citations
- Exploration via Planning for Information about the Optimal TrajectoryViraj Mehta, Ian Char, Joseph Abbate, Rory Conlin et al.NeurIPS 2022 · 12 citations
- Episodic Future Thinking Mechanism for Multi-agent Reinforcement LearningDongsu Lee, Minhae KwonNeurIPS 2024 · 9 citations
- PID-Inspired Inductive Biases for Deep Reinforcement Learning in Partially Observable Control TasksIan Char, Jeff SchneiderNeurIPS 2023 · 8 citations
- DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under UncertaintyMingxuan Cui, Duo Zhou, Yuxuan Han, Grani A. Hanasusanto et al.ICLR 2026 · 6 citations
Builds on7
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky et al.ICML 2020 · 186 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative ModelGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu et al.NeurIPS 2020 · 159 citations
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski et al.ICML 2020 · 52 citations
- Bayesian Algorithm Execution: Estimating Computable Properties of Black-box Functions Using Mutual InformationWillie Neiswanger, Ke Alexander Wang, Stefano ErmonICML 2021 · 40 citations
Related papers
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr et al.ICLR 2020 · 104 citations
- Information Directed Reward Learning for Reinforcement LearningDavid Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek et al.NeurIPS 2021 · 27 citations
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 67 citations
- ALINE: Joint Amortization for Bayesian Inference and Active Data AcquisitionDaolang Huang, Xinyi Wen, Ayush Bharti, Samuel Kaski et al.NeurIPS 2025 · 8 citations
- Transition Constrained Bayesian Optimization via Markov Decision ProcessesJose Pablo Folch, Calvin Tsay, Robert M. Lee, Behrang Shafei et al.NeurIPS 2024 · 10 citations
