High-Dimensional Contextual Policy Search with Unknown Context Rewards using Bayesian Optimization
Qing Feng, Benjamin Letham, Hongzi Mao, Eytan Bakshy
Abstract
Contextual policies are used in many settings to customize system parameters and actions to the specifics of a particular setting. In some real-world settings, such as randomized controlled trials or A/B tests, it may not be possible to measure policy outcomes at the level of context-we observe only aggregate rewards across a distribution of contexts. This makes policy optimization much more difficult because we must solve a high-dimensional optimization problem over the entire space of contextual policies, for which existing optimization methods are not suitable. We develop effective models that leverage the structure of the search space to enable contextual policy optimization directly from the aggregate rewards using Bayesian optimization. We use a collection of simulation studies to characterize the performance and robustness of the models, and show that our approach of inferring a low-dimensional context embedding performs best. Finally, we show successful contextual policy optimization in a real-world video bitrate policy problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02205a42-5c38-47f4-8d2f-0e046c095e89Cited by top-tier papers5
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Parallel Bayesian Optimization of Multiple Noisy Objectives with Expected Hypervolume ImprovementSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2021 · 276 citations
- Bayesian Optimization with High-Dimensional OutputsWesley J. Maddox, Maximilian Balandat, Andrew Gordon Wilson, Eytan BakshyNeurIPS 2021 · 75 citations
- Personalized Dual-Level Color Grading for 360-degree Images in Virtual RealityLinping Yuan, John J. Dudley, Per Ola Kristensson, Huamin QuIEEE VR 2025 · 3 citations
- Parametric Pareto Set Learning for Expensive Multi-Objective OptimizationJi Cheng, Bo Xue, Qingfu ZhangAAAI 2026 · 1 citation
Builds on2
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Re-Examining Linear Embeddings for High-Dimensional Bayesian OptimizationBenjamin Letham, Roberto Calandra, Akshara Rai, Eytan BakshyNeurIPS 2020 · 152 citations
Related papers
- CO-BED: Information-Theoretic Contextual Optimization via Bayesian Experimental DesignDesi R. Ivanova, Joel Jennings, Tom Rainforth, Cheng Zhang et al.ICML 2023 · 4 citations
- Contextual Causal Bayesian OptimisationVahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar, Arnak S. DalalyanICLR 2026 · 3 citations
- Empirical Likelihood for Contextual BanditsNikos Karampatziakis, John Langford, Paul MineiroNeurIPS 2020 · 11 citations
- Control Variates for Slate Off-Policy EvaluationNikos Vlassis, Ashok Chandrashekar, Fernando Amat Gil, Nathan KallusNeurIPS 2021 · 21 citations
- Local Bayesian optimization via maximizing probability of descentQuan Nguyen, Kaiwen Wu, Jacob R. Gardner, Roman GarnettNeurIPS 2022 · 41 citations
