Direct Preference-Based Evolutionary Multi-Objective Optimization with Dueling Bandits
Tian Huang, Shengbo Wang, Ke Li
摘要
The ultimate goal of multi-objective optimization (MO) is to assist human decisionmakers (DMs) in identifying solutions of interest (SOI) that optimally reconcile multiple objectives according to their preferences. Preference-based evolutionary MO (PBEMO) has emerged as a promising framework that progressively approximates SOI by involving human in the optimization-cum-decision-making process. Yet, current PBEMO approaches are prone to be inefficient and misaligned with the DM's true aspirations, especially when inadvertently exploiting mis-calibrated reward models. This is further exacerbated when considering the stochastic nature of human feedback. This paper proposes a novel framework that navigates MO to SOI by directly leveraging human feedback without being restricted by a predefined reward model nor cumbersome model selection. Specifically, we developed a clustering-based stochastic dueling bandits algorithm that strategically scales well to high-dimensional dueling bandits. The learned preferences are then transformed into a unified probabilistic format that can be readily adapted to prevalent EMO algorithms. This also leads to a principled termination criterion that strategically manages human cognitive loads and computational budget. Experiments on 48 benchmark test problems, including the RNA inverse design and protein structure prediction, fully demonstrate the effectiveness of our proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Neural Evolution Strategy for Black-box Pareto Set LearningChengyu Lu, Zhenhua Li, Xi Lin, Ji Cheng 等NeurIPS 2025
- Preference Is More than Comparisons: Rethinking Dueling Bandits with Augmented Human FeedbackShengbo Wang, Hong Sun, Ke LiAAAI 2026
它引用的顶会 Paper4
- Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity ModelsViktor Bengs, Aadirupa Saha, Eyke HüllermeierICML 2022 · 被引用 32 次
- Versatile Dueling Bandits: Best-of-both World Analyses for Learning from Relative PreferencesAadirupa Saha, Pierre GaillardICML 2022 · 被引用 30 次
- Interactively Learning Preference Constraints in Linear BanditsDavid Lindner, Sebastian Tschiatschek, Katja Hofmann, Andreas KrauseICML 2022 · 被引用 15 次
- Human Preferences as Dueling BanditsXinyi Yan, Chengxi Luo, Charles L. A. Clarke, Nick Craswell 等SIGIR 2022 · 被引用 9 次
相关 Paper
- Preference-based Reinforcement Learning with Finite-Time GuaranteesYichong Xu, Ruosong Wang, Lin F. Yang, Aarti Singh 等NeurIPS 2020 · 被引用 82 次
- Online Clustering of Dueling BanditsZhiyong Wang, Jiahang Sun, Mingze Kong, Jize Xie 等ICML 2025
- Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao Luo, Ke LiICML 2026 · 被引用 3 次
- High-Dimensional Dueling Optimization with Preference EmbeddingYangwenhui Zhang, Hong Qian, Xiang Shu, Aimin ZhouAAAI 2023 · 被引用 4 次
- Preference Optimization on Pareto Sets: On a Theory of Multi-Objective OptimizationAbhishek Roy, Geelon So, Yian MaNeurIPS 2025 · 被引用 12 次
