Pareto Set Learning for Multi-Objective Reinforcement Learning
Erlong Liu, Yu-Chang Wu, Xiaobin Huang, Chengrui Gao, Ren-Jian Wang, Ke Xue, Chao Qian
Abstract
Multi-objective decision-making problems have emerged in numerous real-world scenarios, such as video games, navigation and robotics. Considering the clear advantages of Reinforcement Learning (RL) in optimizing decision-making processes, researchers have delved into the development of Multi-Objective RL (MORL) methods for solving multi-objective decision problems. However, previous methods either cannot obtain the entire Pareto front, or employ only a single policy network for all the preferences over multiple objectives, which may not produce personalized solutions for each preference. To address these limitations, we propose a novel decomposition-based framework for MORL, Pareto Set Learning for MORL (PSL-MORL), that harnesses the generation capability of hypernetwork to produce the parameters of the policy network for each decomposition weight, generating relatively distinct policies for various scalarized subproblems with high efficiency. PSL-MORL is a general framework, which is compatible for any RL algorithm. The theoretical result guarantees the superiority of the model capacity of PSL-MORL and the optimality of the obtained policy network. Through extensive experiments on diverse benchmarks, we demonstrate the effectiveness of PSL-MORL in achieving dense coverage of the Pareto front, significantly outperforming state-of-the-art MORL methods in both the hypervolume and sparsity indicators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e90afce3-9f49-40e9-b14a-1a0e431fdfc0Cited by top-tier papers5
- Ctrl-DNA: Controllable Cell-Type-Specific Regulatory DNA Design via Constrained RLXingyu Chen, Shihao Ma, Runsheng Lin, Jiecong Lin et al.NeurIPS 2025 · 5 citations
- PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward ModelBaijiong Lin, Weisen Jiang, Yuancheng Xu, Hao Chen et al.ICML 2025
- Neural Evolution Strategy for Black-box Pareto Set LearningChengyu Lu, Zhenhua Li, Xi Lin, Ji Cheng et al.NeurIPS 2025
- LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem ExplorationRuiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li et al.AAAI 2026
- Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-TuningDun Yuan, Di Wu, Xue LiuICLR 2026
Builds on10
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- Learning the Pareto Front with HypernetworksAviv Navon, Aviv Shamsian, Ethan Fetaya, Gal ChechikICLR 2021 · 189 citations
- Pareto Set Learning for Expensive Multi-Objective OptimizationXi Lin, Zhiyuan Yang, Xiaoyuan Zhang, Qingfu ZhangNeurIPS 2022 · 119 citations
- Multi-Objective GFlowNetsMoksh Jain, Sharath Chandra Raparthy, Alex Hernández-García, Jarrid Rector-Brooks et al.ICML 2023 · 113 citations
- Pareto Set Learning for Neural Multi-Objective Combinatorial OptimizationXi Lin, Zhiyuan Yang, Qingfu ZhangICLR 2022 · 105 citations
Related papers
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 8 citations
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song et al.ICLR 2025
- PA2D-MORL: Pareto Ascent Directional Decomposition Based Multi-Objective Reinforcement LearningTianmeng Hu, Biao LuoAAAI 2024 · 7 citations
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
