MTRec: Learning to Align with User Preferences via Mental Reward Models
Mengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li, Pengjie Gu, Zhenhua Dong, Ruiming Tang, Yi Cai
摘要
Recommendation models are predominantly trained using implicit user feedback, since explicit feedback is often costly to obtain. However, implicit feedback, such as clicks, does not always reflect users'real preferences. For example, a user might click on a news article because of its attractive headline, but end up feeling uncomfortable after reading the content. In the absence of explicit feedback, such erroneous implicit signals may severely mislead recommender systems. In this paper, we propose MTRec, a novel sequential recommendation framework designed to align with real user preferences by uncovering their internal satisfaction on recommended items. Specifically, we introduce a mental reward model to quantify user satisfaction and propose a distributional inverse reinforcement learning approach to learn it. The learned mental reward model is then used to guide recommendation models to better align with users'real preferences. Our experiments show that MTRec brings significant improvements to a variety of recommendation models. We also deploy MTRec on an industrial short video platform and observe a 7 percent increase in average user viewing time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song 等NeurIPS 2021 · 被引用 271 次
- Curriculum Disentangled Recommendation with Noisy Multi-feedbackHong Chen, Yudong Chen, Xin Wang, Ruobing Xie 等NeurIPS 2021 · 被引用 88 次
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao 等SIGIR 2023 · 被引用 86 次
- Sampler Design for Implicit Feedback Data by Noisy-label Robust LearningWenhui Yu, Zheng QinSIGIR 2020 · 被引用 54 次
相关 Paper
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen 等ICML 2026 · 被引用 2 次
- Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video RecommendationNa Li, Jiaqi Yu, Minzhi Xie, Tiantian He 等SIGIR 2026
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang 等WWW 2022 · 被引用 14 次
- Treatment Effect Estimation for User Interest Exploration on Recommender SystemsJiaju Chen, Wenjie Wang, Chongming Gao, Peng Wu 等SIGIR 2024 · 被引用 8 次
- FeedRec: News Feed Recommendation with Various User FeedbacksChuhan Wu, Fangzhao Wu, Tao Qi, Qi Liu 等WWW 2022 · 被引用 92 次
