Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative Feedback
Han Shao, Lee Cohen, Avrim Blum, Yishay Mansour, Aadirupa Saha, Matthew R. Walter
摘要
In classic reinforcement learning (RL) and decision making problems, policies are evaluated with respect to a scalar reward function, and all optimal policies are the same with regards to their expected return. However, many real-world problems involve balancing multiple, sometimes conflicting, objectives whose relative priority will vary according to the preferences of each user. Consequently, a policy that is optimal for one user might be sub-optimal for another. In this work, we propose a multi-objective decision making framework that accommodates different user preferences over objectives, where preferences are learned via policy comparisons. Our model consists of a Markov decision process with a vector-valued reward function, with each user having an unknown preference vector that expresses the relative importance of each objective. The goal is to efficiently compute a near-optimal policy for a given user. We consider two user feedback models. We first address the case where a user is provided with two policies and returns their preferred policy as feedback. We then move to a different user feedback model, where a user is instead provided with two small weighted sets of representative trajectories and selects the preferred one. In both cases, we suggest an algorithm that finds a nearly optimal policy for the user using a small number of comparison queries.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Tunable LLM-based Proactive Recommendation AgentMingze Wang, Chongming Gao, Wenjie Wang, Yangyang Li 等ACL 2025 · 被引用 4 次
- SEED-SET: Scalable Evolving Experimental Design for System-level Ethical TestingAnjali Parashar, Yingke Li, Eric Yang Yu, Fei Chen 等ICLR 2026 · 被引用 1 次
- Efficient Discovery of Pareto Front for Multi-Objective Reinforcement LearningRuohong Liu, Yuxin Pan, Linjie Xu, Lei Song 等ICLR 2025
它引用的顶会 Paper4
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Dueling Convex OptimizationAadirupa Saha, Tomer Koren, Yishay MansourICML 2021 · 被引用 22 次
- Preference learning along multiple criteria: A game-theoretic perspectiveKush Bhatia, Ashwin Pananjady, Peter L. Bartlett, Anca D. Dragan 等NeurIPS 2020 · 被引用 15 次
- Online Learning with Primary and Secondary LossesAvrim Blum, Han ShaoNeurIPS 2020 · 被引用 1 次
相关 Paper
- Accommodating Picky Customers: Regret Bound and Exploration Complexity for Multi-Objective Reinforcement LearningJingfeng Wu, Vladimir Braverman, Lin YangNeurIPS 2021 · 被引用 11 次
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
- AMOR: Adaptive Character Control through Multi-Objective Reinforcement LearningLucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller 等SIGGRAPH 2025 · 被引用 4 次
- Towards Theoretical Understanding of Sequential Decision Making with Preference FeedbackSimone Drago, Marco Mussi, Alberto Maria MetelliICML 2025
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 被引用 6 次
