Pareto Set Learning for Multi-Objective Reinforcement Learning
Erlong Liu, Yu-Chang Wu, Xiaobin Huang, Chengrui Gao, Ren-Jian Wang, Ke Xue, Chao Qian
摘要
Multi-objective decision-making problems have emerged in numerous real-world scenarios, such as video games, navigation and robotics. Considering the clear advantages of Reinforcement Learning (RL) in optimizing decision-making processes, researchers have delved into the development of Multi-Objective RL (MORL) methods for solving multi-objective decision problems. However, previous methods either cannot obtain the entire Pareto front, or employ only a single policy network for all the preferences over multiple objectives, which may not produce personalized solutions for each preference. To address these limitations, we propose a novel decomposition-based framework for MORL, Pareto Set Learning for MORL (PSL-MORL), that harnesses the generation capability of hypernetwork to produce the parameters of the policy network for each decomposition weight, generating relatively distinct policies for various scalarized subproblems with high efficiency. PSL-MORL is a general framework, which is compatible for any RL algorithm. The theoretical result guarantees the superiority of the model capacity of PSL-MORL and the optimality of the obtained policy network. Through extensive experiments on diverse benchmarks, we demonstrate the effectiveness of PSL-MORL in achieving dense coverage of the Pareto front, significantly outperforming state-of-the-art MORL methods in both the hypervolume and sparsity indicators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RLXingyu Chen, Shihao Ma, Runsheng Lin, Jiecong Lin 等NeurIPS 2025 · 被引用 5 次
- PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward ModelBaijiong Lin, Weisen Jiang, Yuancheng Xu, Hao Chen 等ICML 2025
- Neural Evolution Strategy for Black-box Pareto Set LearningChengyu Lu, Zhenhua Li, Xi Lin, Ji Cheng 等NeurIPS 2025
- LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem ExplorationRuiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li 等AAAI 2026
- Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-TuningDun Yuan, Di Wu, Xue LiuICLR 2026
它引用的顶会 Paper10
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 被引用 189 次
- Pareto Set Learning for Expensive Multi-Objective OptimizationXi Lin, Zhiyuan Yang, Xiaoyuan Zhang, Qingfu ZhangNeurIPS 2022 · 被引用 119 次
- Multi-Objective GFlowNetsMoksh Jain, Sharath Chandra Raparthy, Alex Hernández-García, Jarrid Rector-Brooks 等ICML 2023 · 被引用 113 次
- Pareto Set Learning for Neural Multi-Objective Combinatorial OptimizationXi Lin, Zhiyuan Yang, Qingfu ZhangICLR 2022 · 被引用 105 次
相关 Paper
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 被引用 8 次
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song 等ICLR 2025
- PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao LuoAAAI 2024 · 被引用 7 次
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
