Lune

NeurIPS2023顶会

Distributional Pareto-Optimal Multi-Objective Reinforcement Learning

Xin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian, Masashi Sugiyama, Ashley Llorens

2023年份
46被引次数
10顶会引用

摘要

Multi-objective reinforcement learning (MORL) has been proposed to learn control 1 policies over multiple competing objectives with each possible preference over 2 returns. However, current MORL algorithms fail to account for distributional 3 preferences over the multi-variate returns, which are particularly important in real-4 world scenarios such as autonomous driving. To address this issue, we extend the 5 concept of Pareto-optimality in MORL into distributional Pareto-optimality, which 6 captures the optimality of return distributions, rather than the expectations. Our 7 proposed method, called Distributional Pareto-Optimal Multi-Objective Reinforce-8 ment Learning (DPMORL), is capable of learning distributional Pareto-optimal 9 policies that balance multiple objectives while considering the return uncertainty. 10 We evaluated our method on several benchmark problems and demonstrated its 11 effectiveness in discovering distributional Pareto-optimal policies and satisfying 12 diverse distributional preferences compared to existing MORL methods.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper10

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖