Distributional Pareto-Optimal Multi-Objective Reinforcement Learning
Xin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian, Masashi Sugiyama, Ashley Llorens
摘要
Multi-objective reinforcement learning (MORL) has been proposed to learn control 1 policies over multiple competing objectives with each possible preference over 2 returns. However, current MORL algorithms fail to account for distributional 3 preferences over the multi-variate returns, which are particularly important in real-4 world scenarios such as autonomous driving. To address this issue, we extend the 5 concept of Pareto-optimality in MORL into distributional Pareto-optimality, which 6 captures the optimality of return distributions, rather than the expectations. Our 7 proposed method, called Distributional Pareto-Optimal Multi-Objective Reinforce-8 ment Learning (DPMORL), is capable of learning distributional Pareto-optimal 9 policies that balance multiple objectives while considering the return uncertainty. 10 We evaluated our method on several benchmark problems and demonstrated its 11 effectiveness in discovering distributional Pareto-optimal policies and satisfying 12 diverse distributional preferences compared to existing MORL methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Foundations of Multivariate Distributional Reinforcement LearningHarley Wiltzer, Jesse Farebrother, Arthur Gretton, Mark RowlandNeurIPS 2024 · 被引用 21 次
- Dual-Objective Reinforcement Learning with Novel Hamilton-Jacobi-Bellman FormulationsWilliam Sharpless, Dylan Hirsch, Sander Tonkens, Nikhil Uday Shinde 等ICLR 2026 · 被引用 12 次
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 被引用 4 次
- FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement LearningWoosung Kim, Jinho Lee, Jongmin Lee, Byung-Jun LeeNeurIPS 2025 · 被引用 4 次
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 被引用 1 次
它引用的顶会 Paper8
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- Learning to Weight Imperfect DemonstrationsYunke Wang, Chang Xu, Bo Du, Honglak LeeICML 2021 · 被引用 57 次
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 被引用 54 次
- Optimistic Linear Support and Successor Features as a Basis for Optimal Policy TransferLucas Nunes Alegre, Ana L. C. Bazzan, Bruno C. da SilvaICML 2022 · 被引用 36 次
- Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsPushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong 等NeurIPS 2021 · 被引用 33 次
相关 Paper
- PD-MORL: Preference-Driven Multi-Objective Reinforcement Learning AlgorithmToygun Basaklar, Suat Gumussoy, Ümit Y. OgrasICLR 2023 · 被引用 8 次
- Preference Controllable Reinforcement Learning with Advanced Multi-Objective OptimizationYucheng Yang, Tianyi Zhou, Mykola Pechenizkiy, Meng FangICML 2025
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song 等ICLR 2025
- Scaling Pareto-Efficient Decision Making via Offline Multi-Objective RLBaiting Zhu, Meihua Dang, Aditya GroverICLR 2023 · 被引用 1 次
- On Generalization Across Environments In Multi-Objective Reinforcement LearningJayden Teoh, Pradeep Varakantham, Peter VamplewICLR 2025
