Policy Search via Bayesian Optimization with Temporal Difference Gaussian Processes
Armin Lederer, Anuj Srivastava, Marco Bagatella, Andreas Krause
摘要
Bayesian optimization (BO) is a method commonly used for policy search in problems with low-dimensional policy parameterizations. While it is generally considered data-efficient, existing BO approaches are agnostic to the sequential structure of the optimization objective induced by policy roll-outs. Thereby, valuable information is discarded that could improve the convergence of BO. We address this inefficiency by developing and rigorously analyzing a novel approach for BO that relies on a temporal difference learning formulation for discounted infinite-horizon value functions based on Gaussian process (GP) regression. We derive learning error bounds for the proposed temporal difference GPs, such that we can exploit upper confidence bounds to analyze the cumulative regret of our BO approach. This analysis is further refined by bounding the maximal information gain for our temporal difference GP model. In a comparison with relevant baseline methods, we demonstrate the practical advantages of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 被引用 308 次
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature MappingDongruo Zhou, Jiafan He, Quanquan GuICML 2021 · 被引用 143 次
- Information Theoretic Regret Bounds for Online Nonlinear ControlSham M. Kakade, Akshay Krishnamurthy, Kendall Lowrey, Motoya Ohnishi 等NeurIPS 2020 · 被引用 137 次
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 被引用 120 次
- Practical and Rigorous Uncertainty Bounds for Gaussian Process RegressionChristian Fiedler, Carsten W. Scherer, Sebastian TrimpeAAAI 2021 · 被引用 92 次
相关 Paper
- On Regret Bounds of Thompson Sampling for Bayesian OptimizationShion Takeno, Shogo IwazakiICML 2026 · 被引用 3 次
- Objective Bound Conditional Gaussian Process for Bayesian OptimizationTaewon Jeong, Heeyoung KimICML 2021 · 被引用 3 次
- Exploring and Exploiting Model Uncertainty in Bayesian OptimizationZishi Zhang, Tao Ren, Yijie PengNeurIPS 2025 · 被引用 1 次
- Bayesian Optimization under Stochastic Delayed FeedbackArun Verma, Zhongxiang Dai, Bryan Kian Hsiang LowICML 2022 · 被引用 15 次
- Trading Convergence Rate with Computational Budget in High Dimensional Bayesian OptimizationHung Tran-The, Sunil Gupta, Santu Rana, Svetha VenkateshAAAI 2020 · 被引用 14 次
