MTRec: Learning to Align with User Preferences via Mental Reward Models
Mengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li, Pengjie Gu, Zhenhua Dong, Ruiming Tang, Yi Cai
Abstract
Recommendation models are predominantly trained using implicit user feedback, since explicit feedback is often costly to obtain. However, implicit feedback, such as clicks, does not always reflect users'real preferences. For example, a user might click on a news article because of its attractive headline, but end up feeling uncomfortable after reading the content. In the absence of explicit feedback, such erroneous implicit signals may severely mislead recommender systems. In this paper, we propose MTRec, a novel sequential recommendation framework designed to align with real user preferences by uncovering their internal satisfaction on recommended items. Specifically, we introduce a mental reward model to quantify user satisfaction and propose a distributional inverse reinforcement learning approach to learn it. The learned mental reward model is then used to guide recommendation models to better align with users'real preferences. Our experiments show that MTRec brings significant improvements to a variety of recommendation models. We also deploy MTRec on an industrial short video platform and observe a 7 percent increase in average user viewing time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a237dbf-e62b-46ce-8d3a-ec3276161d1aBuilds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- IQ-Learn: Inverse soft-Q Learning for ImitationDivyansh Garg, Shuvam Chakraborty, Chris Cundy, Jiaming Song et al.NeurIPS 2021 · 271 citations
- Curriculum Disentangled Recommendation with Noisy Multi-feedbackHong Chen, Yudong Chen, Xin Wang, Ruobing Xie et al.NeurIPS 2021 · 88 citations
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao et al.SIGIR 2023 · 86 citations
- Sampler Design for Implicit Feedback Data by Noisy-label Robust LearningWenhui Yu, Zheng QinSIGIR 2020 · 54 citations
Related papers
- Unbiased Reward Modeling from Implicit Feedback for LLM AlignmentHao Wang, Haocheng Yang, Licheng Pan, Zhichao Chen et al.ICML 2026 · 2 citations
- Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video RecommendationNa Li, Jiaqi Yu, Minzhi Xie, Tiantian He et al.SIGIR 2026
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang et al.WWW 2022 · 14 citations
- Treatment Effect Estimation for User Interest Exploration on Recommender SystemsJiaju Chen, Wenjie Wang, Chongming Gao, Peng Wu et al.SIGIR 2024 · 8 citations
- FeedRec: News Feed Recommendation with Various User FeedbacksChuhan Wu, Fangzhao Wu, Tao Qi, Qi Liu et al.WWW 2022 · 92 citations
