Thompson Sampling for Multi-Objective Linear Contextual Bandit
Somangchan Park, Heesang Ann, Min-hwan Oh
摘要
We study the multi-objective linear contextual bandit problem, where multiple possible conflicting objectives must be optimized simultaneously. We propose MOL-TS, the first Thompson Sampling algorithm with Pareto regret guarantees for this problem. Unlike standard approaches that compute an empirical Pareto front each round, MOL-TS samples parameters across objectives and efficiently selects an arm from a novel effective Pareto front, which accounts for repeated selections over time. Our analysis shows that MOL-TS achieves a worst-case Pareto regret bound of , where is the dimension of the feature vectors, is the total number of rounds, matching the best known order for randomized linear bandit algorithms for single objective. Empirical results confirm the benefits of our proposed approach, demonstrating improved regret minimization and strong multi-objective performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Random Hypervolume Scalarizations for Provable Multi-Objective Black Box OptimizationQiuyi (Richard) Zhang, Daniel GolovinICML 2020 · 被引用 96 次
- Pareto Regret Analyses in Multi-objective Multi-armed BanditMengfan Xu, Diego KlabjanICML 2023 · 被引用 15 次
- Optimal Scalarizations for Sublinear Hypervolume RegretQiuyi (Richard) ZhangNeurIPS 2024 · 被引用 9 次
- Combinatorial Neural BanditsTaehyun Hwang, Kyuwook Chai, Min-hwan OhICML 2023 · 被引用 7 次
- Hierarchize Pareto Dominance in Multi-Objective Stochastic Linear BanditsJi Cheng, Bo Xue, Jiaxiang Yi, Qingfu ZhangAAAI 2024 · 被引用 5 次
相关 Paper
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 被引用 4 次
- Doubly Robust Thompson Sampling with Linear PayoffsWonyoung Kim, Gi-Soo Kim, Myunghee Cho PaikNeurIPS 2021 · 被引用 35 次
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
- Noise-Adaptive Thompson Sampling for Linear Contextual BanditsRuitu Xu, Yifei Min, Tianhao WangNeurIPS 2023 · 被引用 19 次
- Improved Regret of Linear Ensemble SamplingHarin Lee, Min-hwan OhNeurIPS 2024 · 被引用 8 次
