Estimating and Penalizing Induced Preference Shifts in Recommender Systems
Micah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-Menell
Abstract
The content that a recommender system (RS) shows to users influences them. Therefore, when choosing a recommender to deploy, one is implicitly also choosing to induce specific internal states in users. Even more, systems trained via long-horizon optimization will have direct incentives to manipulate users: in this work, we focus on the incentive to shift user preferences so they are easier to satisfy. We argue that – before deployment – system designers should: estimate the shifts a recommender would induce; evaluate whether such shifts would be undesirable; and perhaps even actively optimize to avoid problematic shifts. These steps involve two challenging ingredients: estimation requires anticipating how hypothetical algorithms would influence user preferences if deployed – we do this by using historical user interaction data to train a predictive user model which implicitly contains their preference dynamics; evaluation and optimization additionally require metrics to assess whether such influences are manipulative or otherwise unwanted – we use the notion of “safe shifts”, that define a trust region within which behavior is safe: for instance, the natural way in which users would shift without interference from the system could be deemed “safe”. In simulated experiments, we show that our learned preference dynamics model is effective in estimating user preferences and how they would respond to new recommenders. Additionally, we show that recommenders that optimize for staying in the trust region can avoid manipulative behaviors while still generating engagement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Feedback Loops With Language Models Drive In-Context Reward HackingAlexander Pan, Erik Jones, Meena Jagadeesan, Jacob SteinhardtICML 2024 · 67 citations
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- Path-Specific Objectives for Safer Agent IncentivesSebastian Farquhar, Ryan Carey, Tom EverittAAAI 2022 · 30 citations
- Improved Bayes Risk Can Yield Reduced Social Welfare Under CompetitionMeena Jagadeesan, Michael I. Jordan, Jacob Steinhardt, Nika HaghtalabNeurIPS 2023 · 20 citations
- Performative Recommendation: Diversifying Content via Strategic IncentivesItay Eilat, Nir RosenfeldICML 2023 · 19 citations
Builds on2
Related papers
- Harm Mitigation in Recommender Systems under User Preference DynamicsJerry Chee, Shankar Kalyanaraman, Sindhu Kiranmai Ernala, Udi Weinsberg et al.KDD 2024 · 3 citations
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg et al.ICML 2020 · 81 citations
- Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching ApproachMartin Mladenov, Elliot Creager, Omer Ben-Porat, Kevin Swersky et al.ICML 2020 · 70 citations
- Inference-Time Personalized Safety Control via Paired Difference-in-Means InterventionTran Huynh, Ruoxi JiaICLR 2026
- User-Creator Feature Polarization in Recommender Systems with Dual InfluenceTao Lin, Kun Jin, Andrew Estornell, Xiaoying Zhang et al.NeurIPS 2024 · 6 citations
