Estimating and Penalizing Induced Preference Shifts in Recommender Systems
Micah D. Carroll, Anca D. Dragan, Stuart Russell, Dylan Hadfield-Menell
摘要
The content that a recommender system (RS) shows to users influences them. Therefore, when choosing a recommender to deploy, one is implicitly also choosing to induce specific internal states in users. Even more, systems trained via long-horizon optimization will have direct incentives to manipulate users: in this work, we focus on the incentive to shift user preferences so they are easier to satisfy. We argue that – before deployment – system designers should: estimate the shifts a recommender would induce; evaluate whether such shifts would be undesirable; and perhaps even actively optimize to avoid problematic shifts. These steps involve two challenging ingredients: estimation requires anticipating how hypothetical algorithms would influence user preferences if deployed – we do this by using historical user interaction data to train a predictive user model which implicitly contains their preference dynamics; evaluation and optimization additionally require metrics to assess whether such influences are manipulative or otherwise unwanted – we use the notion of “safe shifts”, that define a trust region within which behavior is safe: for instance, the natural way in which users would shift without interference from the system could be deemed “safe”. In simulated experiments, we show that our learned preference dynamics model is effective in estimating user preferences and how they would respond to new recommenders. Additionally, we show that recommenders that optimize for staying in the trust region can avoid manipulative behaviors while still generating engagement.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Feedback Loops With Language Models Drive In-Context Reward HackingAlexander Pan, Erik Jones, Meena Jagadeesan, Jacob SteinhardtICML 2024 · 被引用 67 次
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell 等ICML 2024 · 被引用 44 次
- Path-Specific Objectives for Safer Agent IncentivesSebastian Farquhar, Ryan Carey, Tom EverittAAAI 2022 · 被引用 30 次
- Improved Bayes Risk Can Yield Reduced Social Welfare Under CompetitionMeena Jagadeesan, Michael I. Jordan, Jacob Steinhardt, Nika HaghtalabNeurIPS 2023 · 被引用 20 次
- Performative Recommendation: Diversifying Content via Strategic IncentivesItay Eilat, Nir RosenfeldICML 2023 · 被引用 19 次
它引用的顶会 Paper2
相关 Paper
- Harm Mitigation in Recommender Systems under User Preference DynamicsJerry Chee, Shankar Kalyanaraman, Sindhu Kiranmai Ernala, Udi Weinsberg 等KDD 2024 · 被引用 3 次
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Optimizing Long-term Social Welfare in Recommender Systems: A Constrained Matching ApproachMartin Mladenov, Elliot Creager, Omer Ben-Porat, Kevin Swersky 等ICML 2020 · 被引用 70 次
- Inference-Time Personalized Safety Control via Paired Difference-in-Means InterventionTran Huynh, Ruoxi JiaICLR 2026
- User-Creator Feature Polarization in Recommender Systems with Dual InfluenceTao Lin, Kun Jin, Andrew Estornell, Xiaoying Zhang 等NeurIPS 2024 · 被引用 6 次
