Accommodating Picky Customers: Regret Bound and Exploration Complexity for Multi-Objective Reinforcement Learning
Jingfeng Wu, Vladimir Braverman, Lin Yang
Abstract
In this paper we consider multi-objective reinforcement learning where the objectives are balanced using preferences. In practice, the preferences are often given in an adversarial manner, e.g., customers can be picky in many applications. We formalize this problem as an episodic learning problem on a Markov decision process, where transitions are unknown and a reward function is the inner product of a preference vector with pre-specified multi-objective reward functions. In the online setting, the agent receives a (adversarial) preference every episode and proposes policies to interact with the environment. We provide a model-based algorithm that achieves a regret bound , where is the number of objectives, is the number of states, is the number of actions, is the length of the horizon, and is the number of episodes. Furthermore, we consider preference-free exploration, i.e., the agent first interacts with the environment without specifying any preference and then is able to accommodate arbitrary preference vectors up to error. Our proposed algorithm is provably efficient with a nearly optimal sample complexity .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c95c70a9-3575-459b-bb3e-1615342f67e1Cited by top-tier papers6
- A Simple Reward-free Approach to Constrained Reinforcement LearningSobhan Miryoosefi, Chi JinICML 2022 · 36 citations
- Provably Efficient Algorithms for Multi-Objective Competitive RLTiancheng Yu, Yi Tian, Jingzhao Zhang, Suvrit SraICML 2021 · 25 citations
- Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement LearningRuida Zhou, Tao Liu, Dileep Kalathil, P. R. Kumar et al.NeurIPS 2022 · 22 citations
- Provably Feedback-Efficient Reinforcement Learning via Active Reward LearningDingwen Kong, Lin YangNeurIPS 2022 · 19 citations
- Improved Sample Complexity for Reward-free Reinforcement Learning under Low-rank MDPsYuan Cheng, Ruiquan Huang, Yingbin Liang, Jing YangICLR 2023
Builds on5
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
- On Reward-Free Reinforcement Learning with Linear Function ApproximationRuosong Wang, Simon S. Du, Lin F. Yang, Ruslan SalakhutdinovNeurIPS 2020 · 121 citations
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann et al.ICML 2021 · 110 citations
- Task-agnostic Exploration in Reinforcement LearningXuezhou Zhang, Yuzhe Ma, Adish SinglaNeurIPS 2020 · 56 citations
Related papers
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour et al.NeurIPS 2023 · 10 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
- Provably Efficient Multi-Objective Bandit Algorithms Under Preference-Centric CustomizationLinfeng Cao, Ming Shi, Ness B. ShroffAAAI 2026 · 2 citations
- Multi-objective Linear Reinforcement Learning with Lexicographic RewardsBo Xue, Dake Bu, Ji Cheng, Yuanyu Wan et al.ICML 2025
- Near-Minimax Multi-Objective RL under Predictable Adversarial Preferences and Preference-Free Exploration in Linear MDPsMingxi Hu, Meiling YuICML 2026
