Augmented Bayesian Policy Search
Mahdi Kallel, Debabrota Basu, Riad Akrour, Carlo D'Eramo
Abstract
Deterministic policies are often preferred over stochastic ones when implemented on physical systems. They can prevent erratic and harmful behaviors while being easier to implement and interpret. However, in practice, exploration is largely performed by stochastic policies. First-order Bayesian Optimization (BO) methods offer a principled way of performing exploration using deterministic policies. This is done through a learned probabilistic model of the objective function and its gradient. Nonetheless, such approaches treat policy search as a black-box problem, and thus, neglect the reinforcement learning nature of the problem. In this work, we leverage the performance difference lemma to introduce a novel mean function for the probabilistic model. This results in augmenting BO methods with the action-value function. Hence, we call our method Augmented Bayesian Search (ABS). Interestingly, this new mean function enhances the posterior gradient with the deterministic policy gradient, effectively bridging the gap between BO and policy gradient methods. The resulting algorithm combines the convenience of the direct policy search with the scalability of reinforcement learning. We validate ABS on high-dimensional locomotion problems and demonstrate competitive performance compared to existing direct policy search schemes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e094569b-3e51-47c0-9c95-8bc4088b17a6Cited by top-tier papers2
- Performative Policy Gradient: Optimality in Performative Reinforcement LearningDebabrota Basu, Udvas Das, Brahim Driss, Uddalak MukherjeeICML 2026 · 2 citations
- Policy Search via Bayesian Optimization with Temporal Difference Gaussian ProcessesArmin Lederer, Anuj Srivastava, Marco Bagatella, Andreas KrauseICML 2026
Builds on6
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Learning Search Space Partition for Black-box Optimization using Monte Carlo Tree SearchLinnan Wang, Rodrigo Fonseca, Yuandong TianNeurIPS 2020 · 163 citations
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 67 citations
- Local Bayesian optimization via maximizing probability of descentQuan Nguyen, Kaiwen Wu, Jacob R. Gardner, Roman GarnettNeurIPS 2022 · 41 citations
- Are Random Decompositions all we need in High Dimensional Bayesian Optimisation?Juliusz Krysztof Ziomek, Haitham Bou-AmmarICML 2023 · 39 citations
Related papers
- Bayesian Optimization over Discrete and Mixed Spaces via Probabilistic ReparameterizationSamuel Daulton, Xingchen Wan, David Eriksson, Maximilian Balandat et al.NeurIPS 2022 · 71 citations
- Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement LearningRuiyao Miao, Junren Xiao, Shiya Tsang, Hui Xiong et al.NeurIPS 2025 · 2 citations
- Re-Examining Linear Embeddings for High-Dimensional Bayesian OptimizationBenjamin Letham, Roberto Calandra, Akshara Rai, Eytan BakshyNeurIPS 2020 · 152 citations
- Batched Energy-Entropy acquisition for Bayesian OptimizationFelix Teufel, Carsten Stahlhut, Jesper Ferkinghoff-BorgNeurIPS 2024 · 3 citations
- Transition Constrained Bayesian Optimization via Markov Decision ProcessesJose Pablo Folch, Calvin Tsay, Robert M. Lee, Behrang Shafei et al.NeurIPS 2024 · 10 citations
