Thompson Sampling for Multi-Objective Linear Contextual Bandit
Somangchan Park, Heesang Ann, Min-hwan Oh
Abstract
We study the multi-objective linear contextual bandit problem, where multiple possible conflicting objectives must be optimized simultaneously. We propose MOL-TS, the first Thompson Sampling algorithm with Pareto regret guarantees for this problem. Unlike standard approaches that compute an empirical Pareto front each round, MOL-TS samples parameters across objectives and efficiently selects an arm from a novel effective Pareto front, which accounts for repeated selections over time. Our analysis shows that MOL-TS achieves a worst-case Pareto regret bound of , where is the dimension of the feature vectors, is the total number of rounds, matching the best known order for randomized linear bandit algorithms for single objective. Empirical results confirm the benefits of our proposed approach, demonstrating improved regret minimization and strong multi-objective performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff468cc4-8d75-4c1e-b212-d4d0bd523978Builds on5
- Random Hypervolume Scalarizations for Provable Multi-Objective Black Box OptimizationQiuyi (Richard) Zhang, Daniel GolovinICML 2020 · 96 citations
- Pareto Regret Analyses in Multi-objective Multi-armed BanditMengfan Xu, Diego KlabjanICML 2023 · 15 citations
- Optimal Scalarizations for Sublinear Hypervolume RegretQiuyi (Richard) ZhangNeurIPS 2024 · 9 citations
- Combinatorial Neural BanditsTaehyun Hwang, Kyuwook Chai, Min-hwan OhICML 2023 · 7 citations
- Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear BanditsJi Cheng, Bo Xue, Jiaxiang Yi, Qingfu ZhangAAAI 2024 · 5 citations
Related papers
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 4 citations
- Doubly Robust Thompson Sampling with Linear PayoffsWonyoung Kim, Gi-Soo Kim, Myunghee Cho PaikNeurIPS 2021 · 35 citations
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
- Noise-Adaptive Thompson Sampling for Linear Contextual BanditsRuitu Xu, Yifei Min, Tianhao WangNeurIPS 2023 · 19 citations
- Improved Regret of Linear Ensemble SamplingHarin Lee, Min-hwan OhNeurIPS 2024 · 8 citations
