Inferring Lexicographically-Ordered Rewards from Preferences
Alihan Hüyük, William R. Zame, Mihaela van der Schaar
摘要
Modeling the preferences of agents over a set of alternatives is a principal concern in many areas. The dominant approach has been to find a single reward/utility function with the property that alternatives yielding higher rewards are preferred over alternatives yielding lower rewards. However, in many settings, preferences are based on multiple—often competing—objectives; a single reward function is not adequate to represent such preferences. This paper proposes a method for inferring multi-objective reward-based representations of an agent's observed preferences. We model the agent's priorities over different objectives as entering lexicographically, so that objectives with lower priorities matter only when the agent is indifferent with respect to objectives with higher priorities. We offer two example applications in healthcare—one inspired by cancer treatment, the other inspired by organ transplantation—to illustrate how the lexicographically-ordered rewards we learn can provide a better understanding of a decision-maker's preferences and help improve policies when used in reinforcement learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 被引用 48 次
- Inverse Contextual Bandits: Learning How Behavior Evolves over TimeAlihan Hüyük, Daniel Jarrett, Mihaela van der SchaarICML 2022 · 被引用 14 次
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 被引用 13 次
- Defining Expertise: Applications to Treatment Effect EstimationAlihan Hüyük, Qiyao Wei, Alicia Curth, Mihaela van der SchaarICLR 2024 · 被引用 3 次
它引用的顶会 Paper1
相关 Paper
- A distributional view on multi-objective policy optimizationAbbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever, Michael Neunert 等ICML 2020 · 被引用 93 次
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour 等NeurIPS 2023 · 被引用 10 次
- Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic BanditsBo Xue, Yuanyu Wan, Zhichao Lu, Qingfu ZhangAAAI 2026
- Multi-objective Linear Reinforcement Learning with Lexicographic RewardsBo Xue, Dake Bu, Ji Cheng, Yuanyu Wan 等ICML 2025
- Radiology Report Generation via Multi-objective Preference OptimizationTing Xiao, Lei Shi, Peng Liu, Zhe Wang 等AAAI 2025 · 被引用 21 次
