High-Dimensional Contextual Policy Search with Unknown Context Rewards using Bayesian Optimization
Qing Feng, Benjamin Letham, Hongzi Mao, Eytan Bakshy
摘要
Contextual policies are used in many settings to customize system parameters and actions to the specifics of a particular setting. In some real-world settings, such as randomized controlled trials or A/B tests, it may not be possible to measure policy outcomes at the level of context-we observe only aggregate rewards across a distribution of contexts. This makes policy optimization much more difficult because we must solve a high-dimensional optimization problem over the entire space of contextual policies, for which existing optimization methods are not suitable. We develop effective models that leverage the structure of the search space to enable contextual policy optimization directly from the aggregate rewards using Bayesian optimization. We use a collection of simulation studies to characterize the performance and robustness of the models, and show that our approach of inferring a low-dimensional context embedding performs best. Finally, we show successful contextual policy optimization in a real-world video bitrate policy problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume ImprovementSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2021 · 被引用 276 次
- Bayesian Optimization with High-Dimensional OutputsWesley J. Maddox, Maximilian Balandat, Andrew Gordon Wilson, Eytan BakshyNeurIPS 2021 · 被引用 75 次
- Personalized Dual-Level Color Grading for 360-degree Images in Virtual RealityLinping Yuan, John J. Dudley, Per Ola Kristensson, Huamin QuIEEE VR 2025 · 被引用 3 次
- Parametric Pareto Set Learning for Expensive Multi-Objective OptimizationJi Cheng, Bo Xue, Qingfu ZhangAAAI 2026 · 被引用 1 次
它引用的顶会 Paper2
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
- Re-Examining Linear Embeddings for High-Dimensional Bayesian OptimizationBenjamin Letham, Roberto Calandra, Akshara Rai, Eytan BakshyNeurIPS 2020 · 被引用 152 次
相关 Paper
- CO-BED: Information-Theoretic Contextual Optimization via Bayesian Experimental DesignDesi R. Ivanova, Joel Jennings, Tom Rainforth, Cheng Zhang 等ICML 2023 · 被引用 4 次
- Contextual Causal Bayesian OptimisationVahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar, Arnak S. DalalyanICLR 2026 · 被引用 3 次
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 被引用 11 次
- Control Variates for Slate Off-Policy EvaluationNikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan KallusNeurIPS 2021 · 被引用 21 次
- Local Bayesian optimization via maximizing probability of descentQuan Nguyen, Kaiwen Wu, Jacob R. Gardner, Roman GarnettNeurIPS 2022 · 被引用 41 次
