Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
Xianwei Cao, Dou Quan, Zhenliang Zhang, Shuang Wang
Abstract
Humans often juggle multiple, sometimes conflicting objectives and shift their priorities as circumstances change, rather than following a fixed objective function. In contrast, most computational decision-making and multi-objective RL methods assume static preference weights or a known scalar reward. In this work, we study sequential decision-making problem when these preference weights are unobserved latent variables that drift with context. Specifically, we propose Dynamic Preference Inference (DPI), a cognitively inspired framework in which an agent maintains a probabilistic belief over preference weights, updates this belief from recent interaction, and conditions its policy on inferred preferences. We instantiate DPI as a variational preference inference module trained jointly with a preference-conditioned actor–critic, using vector-valued returns as evidence about latent trade-offs. In queueing, gridworld maze, and multi-objective continuous-control environments with event-driven changes in objectives, DPI adapts its inferred preferences to new regimes and achieves higher post-shift performance than fixed-weight and heuristic envelope baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3efd61a7-4628-4f2e-9ab4-9f5014a20ce8Builds on2
Related papers
- Eliciting User Preferences for Personalized Multi-Objective Decision Making through Comparative FeedbackHan Shao, Lee Cohen, Avrim Blum, Yishay Mansour et al.NeurIPS 2023 · 10 citations
- Inferring Lexicographically-Ordered Rewards from PreferencesAlihan Hüyük, William R. Zame, Mihaela van der SchaarAAAI 2022 · 6 citations
- Regularized Conditional Diffusion Model for Multi-Task Preference AlignmentXudong Yu, Chenjia Bai, Haoran He, Changhong Wang et al.NeurIPS 2024 · 11 citations
- What Is It You Really Want of Me? Generalized Reward Learning with Biased Beliefs about Domain DynamicsZe Gong, Yu ZhangAAAI 2020 · 14 citations
- AMOR: Adaptive Character Control through Multi-Objective Reinforcement LearningLucas N. Alegre, Agon Serifi, Ruben Grandia, David Müller et al.SIGGRAPH 2025 · 4 citations
