Local policy search with Bayesian optimization
Sarah Müller, Alexander von Rohr, Sebastian Trimpe
摘要
Reinforcement learning (RL) aims to find an optimal policy by interaction with an environment. Consequently, learning complex behavior requires a vast number of samples, which can be prohibitive in practice. Nevertheless, instead of systematically reasoning and actively choosing informative samples, policy gradients for local search are often obtained from random perturbations. These random samples yield high variance estimates and hence are sub-optimal in terms of sample complexity. Actively selecting informative samples is at the core of Bayesian optimization, which constructs a probabilistic surrogate of the objective from past samples to reason about informative subsequent ones. In this paper, we propose to join both worlds. We develop an algorithm utilizing a probabilistic model of the objective function and its gradient. Based on the model, the algorithm decides where to query a noisy zeroth-order oracle to improve the gradient estimates. The resulting algorithm is a novel type of policy search method, which we compare to existing black-box algorithms. The comparison reveals improved sample complexity and reduced variance in extensive empirical evaluations on synthetic objectives. Further, we highlight the benefits of active sampling on popular RL benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Vanilla Bayesian Optimization Performs Great in High DimensionsCarl Hvarfner, Erik Orm Hellsten, Luigi NardiICML 2024 · 被引用 88 次
- Local Bayesian optimization via maximizing probability of descentQuan Nguyen, Kaiwen Wu, Jacob R. Gardner, Roman GarnettNeurIPS 2022 · 被引用 41 次
- The Behavior and Convergence of Local Bayesian OptimizationKaiwen Wu, Kyurae Kim, Roman Garnett, Jacob R. GardnerNeurIPS 2023 · 被引用 27 次
- Minimizing UCB: a Better Local Search Strategy in Local Bayesian OptimizationZheyi Fan, Wenyu Wang, Szu Hui Ng, Qingpei HuNeurIPS 2024 · 被引用 14 次
- User Preference Meets Pareto-Optimality in Multi-Objective Bayesian OptimizationJoshua Hang Sai Ip, Ankush Chakrabarty, Ali Mesbah, Diego RomeresAAAI 2025 · 被引用 8 次
它引用的顶会 Paper3
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree SearchLinnan Wang, Rodrigo Fonseca, Yuandong TianNeurIPS 2020 · 被引用 163 次
- Parameter-Based Value FunctionsFrancesco Faccio, Louis Kirsch, Jürgen SchmidhuberICLR 2021 · 被引用 29 次
相关 Paper
- Augmented Bayesian Policy SearchMahdi Kallel, Debabrota Basu, Riad Akrour, Carlo D'EramoICLR 2024 · 被引用 4 次
- BayeSQP: Bayesian Optimization through Sequential Quadratic ProgrammingPaul Brunzema, Sebastian TrimpeNeurIPS 2025 · 被引用 7 次
- Sample Efficient Deep Reinforcement Learning via Uncertainty EstimationVincent Mai, Kaustubh Mani, Liam PaullICLR 2022 · 被引用 53 次
- Making RL with Preference-based Feedback Efficient via RandomizationRunzhe Wu, Wen SunICLR 2024 · 被引用 44 次
- Breaking the Computational Barrier: Provably Efficient Actor–Critic for Low-Rank MDPsRuiquan Huang, Donghao Li, Yingbin LIANG, Jing YangICML 2026
