Multi-objective Linear Reinforcement Learning with Lexicographic Rewards
Bo Xue, Dake Bu, Ji Cheng, Yuanyu Wan, Qingfu Zhang
摘要
Reinforcement Learning (RL) with linear transition kernels and reward functions has recently attracted growing attention due to its computational efficiency and theoretical advancements. However, prior theoretical research in RL has primarily focused on single-objective problems, resulting in limited theoretical development for multi-objective reinforcement learning (MORL). To bridge this gap, we examine MORL under lexicographic reward structures, where rewards comprise m hierarchically ordered objectives. In this framework, the agent maximizes objectives sequentially, prioritizing the highest-priority objective before considering subsequent ones. We introduce the first MORL algorithm with provable regret guarantees. For any objective i ∈ 1, 2, . . . , m, our algorithm achieves a regret bound of O(Λ i (λ) , λ quantifies the trade-off between conflicting objectives, d is the feature dimension, H is the episode length, and K is the number of episodes. Furthermore, our algorithm can be applied in the misspecified setting, where the regret bound for the i-th objective becomes O(Λ i (λ) • ( √ d 2 H 4 K + dH 2 K)), with denoting the degree of misspecification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang 等ICLR 2026 · 被引用 4 次
- Offline Multi-Objective Bandits: From Logged Data to Pareto-Optimal PoliciesJi Cheng, Song Lai, Shunyu Yao, Bo XueAAAI 2026 · 被引用 1 次
- Near-Minimax Multi-Objective RL under Predictable Adversarial Preferences and Preference-Free Exploration in Linear MDPsMingxi Hu, Meiling YuICML 2026
它引用的顶会 Paper12
- Prediction-Guided Multi-Objective Reinforcement Learning for Continuous Robot ControlJie Xu, Yunsheng Tian, Pingchuan Ma, Daniela Rus 等ICML 2020 · 被引用 210 次
- Logarithmic Regret for Reinforcement Learning with Linear Function ApproximationJiafan He, Dongruo Zhou, Quanquan GuICML 2021 · 被引用 108 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
- Nearly Minimax Optimal Reinforcement Learning for Linear Markov Decision ProcessesJiafan He, Heyang Zhao, Dongruo Zhou, Quanquan GuICML 2023 · 被引用 68 次
- A Multi-objective / Multi-task Learning Framework Induced by Pareto StationarityMichinari Momma, Chaosheng Dong, Jia LiuICML 2022 · 被引用 62 次
相关 Paper
- Multiple Trade-offs: An Improved Approach for Lexicographic Linear BanditsBo Xue, Xi Lin, Xiaoyuan Zhang, Qingfu ZhangAAAI 2025 · 被引用 4 次
- Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic BanditsBo Xue, Yuanyu Wan, Zhichao Lu, Qingfu ZhangAAAI 2026
- Multiobjective Lipschitz Bandits under Lexicographic OrderingBo Xue, Ji Cheng, Fei Liu, Yimu Wang 等AAAI 2024 · 被引用 4 次
- LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem ExplorationRuiyu Qiu, Rui Wang, Guanghui Yang, Xiang Li 等AAAI 2026
- An Analytical Study of Utility Functions in Multi-Objective Reinforcement LearningManel Rodriguez-Soto, Juan A. Rodríguez-Aguilar, Maite López-SánchezNeurIPS 2024 · 被引用 9 次
