Local policy search with Bayesian optimization
Sarah Müller, Alexander von Rohr, Sebastian Trimpe
Abstract
Reinforcement learning (RL) aims to find an optimal policy by interaction with an environment. Consequently, learning complex behavior requires a vast number of samples, which can be prohibitive in practice. Nevertheless, instead of systematically reasoning and actively choosing informative samples, policy gradients for local search are often obtained from random perturbations. These random samples yield high variance estimates and hence are sub-optimal in terms of sample complexity. Actively selecting informative samples is at the core of Bayesian optimization, which constructs a probabilistic surrogate of the objective from past samples to reason about informative subsequent ones. In this paper, we propose to join both worlds. We develop an algorithm utilizing a probabilistic model of the objective function and its gradient. Based on the model, the algorithm decides where to query a noisy zeroth-order oracle to improve the gradient estimates. The resulting algorithm is a novel type of policy search method, which we compare to existing black-box algorithms. The comparison reveals improved sample complexity and reduced variance in extensive empirical evaluations on synthetic objectives. Further, we highlight the benefits of active sampling on popular RL benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Vanilla Bayesian Optimization Performs Great in High DimensionsCarl Hvarfner, Erik Orm Hellsten, Luigi NardiICML 2024 · 88 citations
- Local Bayesian optimization via maximizing probability of descentQuan Nguyen, Kaiwen Wu, Jacob R. Gardner, Roman GarnettNeurIPS 2022 · 41 citations
- The Behavior and Convergence of Local Bayesian OptimizationKaiwen Wu, Kyurae Kim, Roman Garnett, Jacob R. GardnerNeurIPS 2023 · 27 citations
- Minimizing UCB: a Better Local Search Strategy in Local Bayesian OptimizationZheyi Fan, Wenyu Wang, Szu Hui Ng, Qingpei HuNeurIPS 2024 · 14 citations
- User Preference Meets Pareto-Optimality in Multi-Objective Bayesian OptimizationJoshua Hang Sai Ip, Ankush Chakrabarty, Ali Mesbah, Diego RomeresAAAI 2025 · 8 citations
Builds on3
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree SearchLinnan Wang, Rodrigo Fonseca, Yuandong TianNeurIPS 2020 · 163 citations
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 29 citations
Related papers
- Augmented Bayesian Policy SearchMahdi Kallel, Debabrota Basu, Riad Akrour, Carlo D'EramoICLR 2024 · 4 citations
- BayeSQP: Bayesian Optimization through Sequential Quadratic ProgrammingPaul Brunzema, Sebastian TrimpeNeurIPS 2025 · 7 citations
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 53 citations
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 44 citations
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
