The Burden of Interactive Alignment with Inconsistent Preferences
Ali Shirali
Abstract
From media platforms to chatbots, algorithms shape how people interact, learn, and discover information. Such interactions between users and an algorithm often unfold over multiple steps, during which strategic users can guide the algorithm to better align with their true interests by selectively engaging with content. However, users frequently exhibit inconsistent preferences: they may spend considerable time on content that offers little long-term value, inadvertently signaling that such content is desirable. Focusing on the user side, this raises a key question: what does it take for such users to align the algorithm with their true interests? To investigate these dynamics, we model the user's decision process as split between a rational "system 2" that decides whether to engage and an impulsive "system 1" that determines how long engagement lasts. We then study a multi-leader, singlefollower extensive Stackelberg game, where users, specifically system 2, lead by committing to engagement strategies and the algorithm best-responds based on observed interactions. We define the burden of alignment as the minimum horizon over which users must optimize to effectively steer the algorithm. We show that a critical horizon exists: users who are sufficiently foresighted can achieve alignment, while those who are not are instead aligned to the algorithm's objective. This critical horizon can be long, imposing a substantial burden. However, even a small, costly signal (e.g., an extra click) can significantly reduce it. Overall, our framework explains how users with inconsistent preferences can align an engagement-driven algorithm with their interests in a Stackelberg equilibrium, highlighting both the challenges and potential remedies for achieving alignment.
Our setting and model. When users have consistent preferences-i.e., the engagement length is in proportion to user's true reward-the alignment problem reduces to engagement maximization. Recent advances in designing instruction-following language models rely on this assumption of consistent preferences, often using models such as Bradley-Terry [1] to directly or indirectly infer human rewards and optimize them to maximize human approval [2][3][4]. Similarly, recommender systems are typically designed to optimize recommendations that maximize user engagement [5].
Users can, however, have inconsistent preferences, where their actions may not reflect their true interests. This often occurs when a user's decision results from a combination of impulsive system 1 and rational system 2 processes [6-8], or when consumption choices are influenced by both long-term benefits (enrichment) and a desire for instant gratification (temptation) [9]. In such cases, revealed preferences do not necessarily align with the user's true preferences.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b52af39a-e8cf-40a4-a38f-79ae55a49044Cited by top-tier papers2
- Direct Alignment with Heterogeneous PreferencesAli Shirali, Arash Nasr-Esfahany, Abdullah Omar Alomar, Parsa Mirtaheri et al.NeurIPS 2025 · 26 citations
- Routing, Cascades, and User Choice for LLMsRafid MahmoodICLR 2026 · 2 citations
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Fine-tuning language models to find agreement among humans with diverse preferencesMichiel A. Bakker, Martin J. Chadwick, Hannah Sheahan, Michael Henry Tessler et al.NeurIPS 2022 · 349 citations
- Consequences of Misaligned AISimon Zhuang, Dylan Hadfield-MenellNeurIPS 2020 · 120 citations
- Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted ProgrammingHussein Mozannar, Gagan Bansal, Adam Fourney, Eric HorvitzCHI 2024 · 88 citations
Related papers
- Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game PerspectiveHaichuan Wang, Tao Lin, Lingkai Kong, Ce Li et al.ICML 2026 · 3 citations
- Impact of Decentralized Learning on Player Utilities in Stackelberg GamesKate Donahue, Nicole Immorlica, Meena Jagadeesan, Brendan Lucier et al.ICML 2024 · 9 citations
- Emergent Alignment via CompetitionNatalie Collina, Surbhi Goel, Aaron Roth, Emily Ryu et al.ICML 2026
- Modeling Attrition in Recommender Systems with Departing BanditsOmer Ben-Porat, Lee Cohen, Liu Leqi, Zachary C. Lipton et al.AAAI 2022 · 14 citations
- Impatient Bandits: Optimizing Recommendations for the Long-Term Without DelayThomas M. McDonald, Lucas Maystre, Mounia Lalmas, Daniel Russo et al.KDD 2023 · 12 citations
