Optimal Scalarizations for Sublinear Hypervolume Regret
Qiuyi (Richard) Zhang
Abstract
Scalarization is a general, parallizable technique that can be deployed in any multiobjective setting to reduce multiple objectives into one, yet some have dismissed this versatile approach because linear scalarizations cannot explore concave regions of the Pareto frontier. To that end, we aim to find simple non-linear scalarizations that provably explore a diverse set of objectives on the Pareto frontier, as measured by the dominated hypervolume. We show that hypervolume scalarizations with uniformly random weights achieves an optimal sublinear hypervolume regret bound of , with matching lower bounds that preclude any algorithm from doing better asymptotically. For the setting of multiobjective stochastic linear bandits, we utilize properties of hypervolume scalarizations to derive a novel non-Euclidean analysis to get regret bounds of , removing unnecessary dependencies. We support our theory with strong empirical performance of using non-linear scalarizations that outperforms both their linear counterparts and other standard multiobjective algorithms in a variety of natural settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dd8c97f-34f2-44a3-87e1-5477751f5fc3Cited by top-tier papers3
- Thompson Sampling for Multi-Objective Linear Contextual BanditSomangchan Park, Heesang Ann, Min-hwan OhNeurIPS 2025 · 1 citation
- Near-Minimax Multi-Objective RL under Predictable Adversarial Preferences and Preference-Free Exploration in Linear MDPsMingxi Hu, Meiling YuICML 2026
- MAGO: Beyond Fixed Hyperparameters with Multi-Objective Pareto Optimization for Hybrid LLM ReasoningHongcheng Ding, Xuanze Zhao, Ruiting Deng, Shamsul Nahar Abdullah et al.ICLR 2026
Builds on3
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian OptimizationSamuel Daulton, Maximilian Balandat, Eytan BakshyNeurIPS 2020 · 428 citations
- Random Hypervolume Scalarizations for Provable Multi-Objective Black Box OptimizationQiuyi (Richard) Zhang, Daniel GolovinICML 2020 · 96 citations
Related papers
- Pareto Regret Analyses in Multi-objective Multi-armed BanditMengfan Xu, Diego KlabjanICML 2023 · 15 citations
- Multi-objective optimization via equivariant deep hypervolume approximationJim Boelrijk, Bernd Ensing, Patrick ForréICLR 2023
- Hypervolume Maximization: A Geometric View of Pareto Set LearningXiaoyuan Zhang, Xi Lin, Bo Xue, Yifan Chen et al.NeurIPS 2023 · 40 citations
- Revisiting Scalarization in Multi-Task Learning: A Theoretical PerspectiveYuzheng Hu, Ruicheng Xian, Qilong Wu, Qiuling Fan et al.NeurIPS 2023 · 76 citations
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 4 citations
