Provably Efficient Multi-Objective Bandit Algorithms Under Preference-Centric Customization
Linfeng Cao, Ming Shi, Ness B. Shroff
Abstract
Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences across objectives, resulting in a Pareto-optimal arm that may score high for one user but perform quite poorly for another. This highlights the need for customized learning, a factor often overlooked in prior research. To address this, we study a preference-aware MO-MAB framework in the presence of explicit user preference. It shifts the focus from achieving Pareto optimality to further optimizing within the Pareto front under preference-centric customization. To our knowledge, this is the first theoretical study of customized MO-MAB optimization with explicit user preferences. Motivated by practical applications, we explore two scenarios: unknown preference and hidden preference, each presenting unique challenges for algorithm design and analysis. At the core of our algorithms are preference estimation and preference-aware optimization mechanisms to adapt to user preferences effectively. We further develop novel analytical techniques to establish near-optimal regret of the proposed algorithms. Strong empirical performance confirm the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d180a891-3341-4f0f-a92d-ad6086bf6a9bBuilds on4
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 66 citations
- Personalized Approximate Pareto-Efficient RecommendationRuobing Xie, Yanlei Liu, Shaoliang Zhang, Rui Wang et al.WWW 2021 · 46 citations
- Pareto Regret Analyses in Multi-objective Multi-armed BanditMengfan Xu, Diego KlabjanICML 2023 · 15 citations
- Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear BanditsJi Cheng, Bo Xue, Jiaxiang Yi, Qingfu ZhangAAAI 2024 · 5 citations
Related papers
- Accommodating Picky Customers: Regret Bound and Exploration Complexity for Multi-Objective Reinforcement LearningJingfeng Wu, Vladimir Braverman, Lin YangNeurIPS 2021 · 11 citations
- User Preference Meets Pareto-Optimality in Multi-Objective Bayesian OptimizationJoshua Hang Sai Ip, Ankush Chakrabarty, Ali Mesbah, Diego RomeresAAAI 2025 · 8 citations
- Local Clustering in Contextual Multi-Armed BanditsYikun Ban, Jingrui HeWWW 2021 · 51 citations
- On Regret with Multiple Best ArmsYinglun Zhu, Robert NowakNeurIPS 2020 · 21 citations
- Preference-Guided Multi-Objective UI AdaptationYao Song, Christoph Gebhardt, Yi-Chi Liao, Christian HolzUIST 2025 · 3 citations
