Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative Feedback
Han Shao, Lee Cohen, Avrim Blum, Yishay Mansour, Aadirupa Saha, Matthew R. Walter
Abstract
In classic reinforcement learning (RL) and decision making problems, policies are evaluated with respect to a scalar reward function, and all optimal policies are the same with regards to their expected return. However, many real-world problems involve balancing multiple, sometimes conflicting, objectives whose relative priority will vary according to the preferences of each user. Consequently, a policy that is optimal for one user might be sub-optimal for another. In this work, we propose a multi-objective decision making framework that accommodates different user preferences over objectives, where preferences are learned via policy comparisons. Our model consists of a Markov decision process with a vector-valued reward function, with each user having an unknown preference vector that expresses the relative importance of each objective. The goal is to efficiently compute a near-optimal policy for a given user. We consider two user feedback models. We first address the case where a user is provided with two policies and returns their preferred policy as feedback. We then move to a different user feedback model, where a user is instead provided with two small weighted sets of representative trajectories and selects the preferred one. In both cases, we suggest an algorithm that finds a nearly optimal policy for the user using a small number of comparison queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d1def34b-bb0c-41b6-8961-d4ba6b46391aCited by top-tier papers3
- Tunable LLM-based Proactive Recommendation AgentMingze Wang, Chongming Gao, Wenjie Wang, Yangyang Li et al.ACL 2025 · 4 citations
- SEED-SET: Scalable Evolving Experimental Design for System-level Ethical TestingAnjali Parashar, Yingke Li, Eric Yang Yu, Fei Chen et al.ICLR 2026 · 1 citation
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song et al.ICLR 2025
Builds on4
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Dueling Convex OptimizationAadirupa Saha, Tomer Koren, Yishay MansourICML 2021 · 22 citations
- Preference learning along multiple criteria: A game-theoretic perspectiveKush Bhatia, Ashwin Pananjady, Peter L. Bartlett, Anca D. Dragan et al.NeurIPS 2020 · 15 citations
- Online Learning with Primary and Secondary LossesAvrim Blum, Han ShaoNeurIPS 2020 · 1 citation
Related papers
- Accommodating Picky Customers: Regret Bound and Exploration Complexity for Multi-Objective Reinforcement LearningJingfeng Wu, Vladimir Braverman, Lin YangNeurIPS 2021 · 11 citations
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert et al.ICML 2020 · 93 citations
- AMOR: Adaptive Character Control through Multi-Objective Reinforcement LearningLucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller et al.SIGGRAPH 2025 · 4 citations
- Towards Theoretical Understanding of Sequential Decision Making with Preference FeedbackSimone Drago, Marco Mussi, Alberto Maria MetelliICML 2025
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 6 citations
