Distributional Pareto-Optimal Multi-Objective Reinforcement Learning
Xin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian, Masashi Sugiyama, Ashley Llorens
Abstract
Multi-objective reinforcement learning (MORL) has been proposed to learn control 1 policies over multiple competing objectives with each possible preference over 2 returns. However, current MORL algorithms fail to account for distributional 3 preferences over the multi-variate returns, which are particularly important in real-4 world scenarios such as autonomous driving. To address this issue, we extend the 5 concept of Pareto-optimality in MORL into distributional Pareto-optimality, which 6 captures the optimality of return distributions, rather than the expectations. Our 7 proposed method, called Distributional Pareto-Optimal Multi-Objective Reinforce-8 ment Learning (DPMORL), is capable of learning distributional Pareto-optimal 9 policies that balance multiple objectives while considering the return uncertainty. 10 We evaluated our method on several benchmark problems and demonstrated its 11 effectiveness in discovering distributional Pareto-optimal policies and satisfying 12 diverse distributional preferences compared to existing MORL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3358ffe1-3f11-41e5-b8ea-0c0bad54d23dCited by top-tier papers10
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 21 citations
- Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman FormulationsWilliam Sharpless, Dylan Hirsch, Sander Tonkens, Nikhil Uday Shinde et al.ICLR 2026 · 12 citations
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 4 citations
- FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement LearningWoosung Kim, Jinho Lee, Jongmin Lee, Byung-Jun LeeNeurIPS 2025 · 4 citations
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 1 citation
Builds on8
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus et al.ICML 2020 · 210 citations
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 57 citations
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 54 citations
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 36 citations
- Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsPushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong et al.NeurIPS 2021 · 33 citations
Related papers
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 8 citations
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song et al.ICLR 2025
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 1 citation
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
