Think before Recommendation: Autonomous Reasoning-enhanced Recommender
Xiaoyu Kong, Junguang Jiang, Bin Liu, Ziru Xu, Zhu Han, Jian Xu, Bo Zheng, Jiancan Wu, Xiang Wang
摘要
The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing distillation-based methods suffer from limitations such as the teacher model's insufficient recommendation capability, costly and static supervision, and superficial transfer of reasoning ability. To address these issues, this paper proposes RecZero, a reinforcement learning (RL)-based recommendation paradigm that abandons the traditional multi-model and multi-stage distillation approach. Instead, RecZero trains a single LLM through pure RL to autonomously develop reasoning capabilities for rating prediction. RecZero consists of two key components: (1) "Think-before-Recommendation" prompt construction, which employs a structured reasoning template to guide the model in step-wise analysis of user interests, item features, and user-item compatibility; and (2) rule-based reward modeling, which adopts group relative policy optimization (GRPO) to compute rewards for reasoning trajectories and optimize the LLM. Additionally, the paper explores a hybrid paradigm, RecOne, which combines supervised fine-tuning with RL, initializing the model with coldstart reasoning samples and further optimizing it with RL. Experimental results demonstrate that RecZero and RecOne significantly outperform existing baseline methods on multiple benchmark datasets, validating the superiority of the RL paradigm in achieving autonomous reasoning-enhanced recommender systems. Our codes are available at https://github.com/AkaliKong/RecZero.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reasoning over Semantic IDs Enhances Generative RecommendationYingzhi He, Yan Sun, Junfei Tan, Yuxin Chen 等KDD 2026 · 被引用 15 次
- Uncertainty-aware Generative RecommendationChenxiao Fan, Chongming Gao, Yaxin Gong, Haoyan Liu 等KDD 2026 · 被引用 2 次
- Factorized Latent Reasoning for LLM-based RecommendationTianqi Gao, Chengkai Huang, Zihan Wang, Cao Liu 等SIGIR 2026
- PARIF: Pushing the Pareto Frontier of Instruction Following and Reasoning with Curriculum Reinforcement LearningRongchuan Mu, Zexin Wang, Qianyu Wang, Minghua Ma 等ACL 2026
它引用的顶会 Paper11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- U-BERT: Pre-training User Representations for Improved RecommendationZhaopeng Qiu, Xian Wu, Jingyue Gao, Wei FanAAAI 2021 · 被引用 171 次
- A Review-aware Graph Contrastive Learning Framework for RecommendationJie Shuai, Kun Zhang, Le Wu, Peijie Sun 等SIGIR 2022 · 被引用 170 次
相关 Paper
- Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationJieyong Kim, Hyunseo Kim, Hyunjin Cho, SeongKu Kang 等SIGIR 2025 · 被引用 13 次
- Absolute Zero: Reinforced Self-play Reasoning with Zero DataAndrew Zhao, Yiran Wu, Tong Wu, Quentin Xu 等NeurIPS 2025 · 被引用 361 次
- Reinforced Latent Reasoning for LLM-based RecommendationYang Zhang, Wenxin Xu, Xiaoyan Zhao, Wenjie Wang 等ICLR 2026 · 被引用 67 次
- Rec: Towards Large Recommender Models with ReasoningRunyang You, Yongqi Li, Xinyu Lin, Xin Zhang 等NeurIPS 2025 · 被引用 3 次
- From Prediction to Understanding: Leveraging Reasoning in Large Language Model-based RecommendationsZhi-Yuan Chen, Siyu Lu, Qiang Liu, Xingxing Wang 等WWW 2026
