Principled Preferential Bayesian Optimization
Wenjie Xu, Wenbin Wang, Yuning Jiang, Bratislav Svetozarevic, Colin N. Jones
Abstract
We study the problem of preferential Bayesian optimization (BO), where we aim to optimize a black-box function with only preference feedback over a pair of candidate solutions. Inspired by the likelihood ratio idea, we construct a confidence set of the black-box function using only the preference feedback. An optimistic algorithm with an efficient computational method is then developed to solve the problem, which enjoys an information-theoretic bound on the total cumulative regret, a first-of-its-kind for preferential BO. This bound further allows us to design a scheme to report an estimated best solution, with a guaranteed convergence rate. Experimental results on sampled instances from Gaussian processes, standard test functions, and a thermal comfort optimization problem all show that our method stably achieves better or competitive performance as compared to the existing state-of-the-art heuristics, which, however, do not have theoretical guarantees on regret bounds or convergence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff320371-57c2-4911-84af-a71f174500d3Cited by top-tier papers9
- Optimal Design for Human Preference ElicitationSubhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari, Aniket Deshmukh et al.NeurIPS 2024 · 20 citations
- Bandits with Preference Feedback: A Stackelberg Game PerspectiveBarna Pásztor, Parnian Kassraie, Andreas KrauseNeurIPS 2024 · 12 citations
- Principled Bayesian Optimization in Collaboration with Human ExpertsWenjie Xu, Masaki Adachi, Colin N. Jones, Michael A. OsborneNeurIPS 2024 · 10 citations
- Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement LearningRuiyao Miao, Junren Xiao, Shiya Tsang, Hui Xiong et al.NeurIPS 2025 · 2 citations
- Human-LLM Collaborative Feature Engineering for Tabular DataZhuoyan Li, Aditya Bansal, Jinzhao Li, Shishuang He et al.ICLR 2026 · 2 citations
Builds on13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Constrained Efficient Global Optimization of Expensive Black-box FunctionsWenjie Xu, Yuning Jiang, Bratislav Svetozarevic, Colin N. JonesICML 2023 · 1,916 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and PlanningSebastian Curi, Felix Berkenkamp, Andreas KrauseNeurIPS 2020 · 120 citations
- Sequential gallery for interactive visual design optimizationYuki Koyama, Issei Sato, Masataka GotoSIGGRAPH 2020 · 90 citations
Related papers
- Batch Bayesian optimisation via density-ratio estimation with guaranteesRafael Oliveira, Louis C. Tiao, Fabio T. RamosNeurIPS 2022 · 11 citations
- Projective Preferential Bayesian OptimizationPetrus Mikkola, Milica Todorovic, Jari Järvi, Patrick Rinke et al.ICML 2020 · 24 citations
- On Regret Bounds of Thompson Sampling for Bayesian OptimizationShion Takeno, Shogo IwazakiICML 2026 · 3 citations
- Multi-Objective Bayesian Optimization with Active Preference LearningRyota Ozaki, Kazuki Ishikawa, Youhei Kanzaki, Shion Takeno et al.AAAI 2024 · 18 citations
- Optimizing Conditional Value-At-Risk of Black-Box FunctionsQuoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2021 · 25 citations
