Lune

ICML2025Top-tier venue

Multi-objective Linear Reinforcement Learning with Lexicographic Rewards

Bo Xue, Dake Bu, Ji Cheng, Yuanyu Wan, Qingfu Zhang

2025Year
3Top-tier citations

Abstract

Reinforcement Learning (RL) with linear transition kernels and reward functions has recently attracted growing attention due to its computational efficiency and theoretical advancements. However, prior theoretical research in RL has primarily focused on single-objective problems, resulting in limited theoretical development for multi-objective reinforcement learning (MORL). To bridge this gap, we examine MORL under lexicographic reward structures, where rewards comprise m hierarchically ordered objectives. In this framework, the agent maximizes objectives sequentially, prioritizing the highest-priority objective before considering subsequent ones. We introduce the first MORL algorithm with provable regret guarantees. For any objective i ∈ 1, 2, . . . , m, our algorithm achieves a regret bound of O(Λ i (λ) , λ quantifies the trade-off between conflicting objectives, d is the feature dimension, H is the episode length, and K is the number of episodes. Furthermore, our algorithm can be applied in the misspecified setting, where the regret bound for the i-th objective becomes O(Λ i (λ) • ( √ d 2 H 4 K + dH 2 K)), with denoting the degree of misspecification.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4ed631e8-e857-4f29-9754-eca1cf46b689

Cited by top-tier papers3

Ask how each one uses it

Builds on12

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines