Towards Theoretical Understanding of Sequential Decision Making with Preference Feedback
Simone Drago, Marco Mussi, Alberto Maria Metelli
Abstract
The success of sequential decision-making approaches, such as reinforcement learning (RL), is closely tied to the availability of a reward feedback. However, designing a reward function that encodes the desired objective is a challenging task. In this work, we address a more realistic scenario: sequential decision making with preference feedback provided, for instance, by a human expert. We aim to build a theoretical basis linking preferences, (non-Markovian) utilities, and (Markovian) rewards, and we study the connections between them. First, we model preference feedback using a partial (pre)order over trajectories, enabling the presence of incomparabilities that are common when preferences are provided by humans but are surprisingly overlooked in existing works. Second, to provide a theoretical justification for a common practice, we investigate how a preference relation can be approximated by a multi-objective utility. We introduce a notion of preference-utility compatibility and analyze the computational complexity of this transformation, showing that constructing the minimumdimensional utility is NP-hard. Third, we propose a novel concept of preference-based policy dominance that does not rely on utilities or rewards and discuss the computational complexity of assessing it. Fourth, we develop a computationally efficient algorithm to approximate a utility using (Markovian) rewards and quantify the error in terms of the suboptimality of the optimal policy induced by the approximating reward. This work aims to lay the foundation for a principled approach to sequential decision making from preference feedback, with promising potential applications in RL from human feedback.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2dd8002-d50c-4fc0-abd7-a3566acd535eCited by top-tier papers1
Ask how each one uses itBuilds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Effects of Reward Misspecification: Mapping and Mitigating Misaligned ModelsAlexander Pan, Kush Bhatia, Jacob SteinhardtICLR 2022 · 293 citations
- Guiding Pretraining in Reinforcement Learning with Large Language ModelsYuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas et al.ICML 2023 · 257 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
Related papers
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour et al.NeurIPS 2023 · 10 citations
- Comparing Comparisons: Informative and Easy Human Feedback with Distinguishability QueriesXuening Feng, Zhaohui Jiang, Timo Kaufmann, Eyke Hüllermeier et al.ICML 2025
- Preference Transformer: Modeling Human Preferences using Transformers for RLChangyeon Kim, Jongjin Park, Jinwoo Shin, Honglak Lee et al.ICLR 2023 · 4 citations
- Sequential Preference Ranking for Efficient Reinforcement Learning from Human FeedbackMinyoung Hwang, Gunmin Lee, Hogun Kee, Chanwoo Kim et al.NeurIPS 2023 · 24 citations
- Listwise Reward Estimation for Offline Preference-based Reinforcement LearningHeewoong Choi, Sangwon Jung, Hongjoon Ahn, Taesup MoonICML 2024 · 14 citations
