Lune

ICML2025顶会

Multi-objective Linear Reinforcement Learning with Lexicographic Rewards

Bo Xue, Dake Bu, Ji Cheng, Yuanyu Wan, Qingfu Zhang

出版方
2025年份
3顶会引用

摘要

Reinforcement Learning (RL) with linear transition kernels and reward functions has recently attracted growing attention due to its computational efficiency and theoretical advancements. However, prior theoretical research in RL has primarily focused on single-objective problems, resulting in limited theoretical development for multi-objective reinforcement learning (MORL). To bridge this gap, we examine MORL under lexicographic reward structures, where rewards comprise m hierarchically ordered objectives. In this framework, the agent maximizes objectives sequentially, prioritizing the highest-priority objective before considering subsequent ones. We introduce the first MORL algorithm with provable regret guarantees. For any objective i ∈ 1, 2, . . . , m, our algorithm achieves a regret bound of O(Λ i (λ) , λ quantifies the trade-off between conflicting objectives, d is the feature dimension, H is the episode length, and K is the number of episodes. Furthermore, our algorithm can be applied in the misspecified setting, where the regret bound for the i-th objective becomes O(Λ i (λ) • ( √ d 2 H 4 K + dH 2 K)), with denoting the degree of misspecification.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 4ed631e8-e857-4f29-9754-eca1cf46b689

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖