Think Wise, Collaborate Effectively: A Rationale-Aware LLM-Based Recommender with Reinforcement Learning from Collaborative Signals
Chung Park, Taesan Kim, Hyeongjun Yun, Dongjoon Hong, Junui Hong, Kijung Park, Mincheol Cho, Minsung Choi, Jihwan Seok, Jaegul Choo
Abstract
Large Language Models (LLMs) have recently emerged as powerful reasoning engines in recommender systems, generating natural-language explanations that foster user engagement. However, their recommendation performance remains limited, as they lack exposure to collaborative useritem interaction patterns. In contrast, collaborative filtering (CF) models achieve strong performance by learning from these behavioral patterns at scale. To unify the strengths of both paradigms, we propose TWiCE-Rec (Think Wise, Collaborate Effectively), a rationale-aware LLM-based recommender that incorporates collaborative user-item interactions. In the first stage, we construct a rationale dataset by applying in-context learning with self-annotated curation. A state-of-the-art LLM is guided to generate persuasive rationales that explain the causal relationship between the user's interaction sequence and the ground-truth next item, resulting in a curated post-hoc training dataset. In the second stage, we perform multi-task instructiontuned adaptation-based on the rationale-augmented training dataset-comprising item description generation and both non-reasoning and reasoning-based sequential recommendation, to equip the LLM with the ability to generate rationales that reflect how user preferences align with item characteristics. Finally, we aim to enhance the LLM's recommendation performance by incorporating user-item interaction patterns derived from the CF-Rec model. To achieve this, we propose a confidence-weighted reinforcement learning strategy that adjusts rewards in proportion to both the LLM's prediction alignment with the ground-truth and the confidence from the pretrained CF-Rec model. Our method outperforms both CFand LLM-Rec models on Amazon datasets in terms of recommendation performance and rationale quality. In an online A/B test, it achieved about 8% higher click-through rate than existing models, demonstrating practical value.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic EncodersYupeng Hou, Jiacheng Li, Xiangjun Fu, Zhankui He et al.ACL 2026 · 346 citations
- RecExplainer: Aligning Large Language Models for Explaining Recommendation ModelsYuxuan Lei, Jianxun Lian, Jing Yao, Xu Huang et al.KDD 2024 · 18 citations
Related papers
- ThinkRec: Thinking-based recommendation via LLMQihang Yu, Kairui Fu, Zheqi Lv, Shengyu Zhang et al.WWW 2026 · 10 citations
- LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential RecommendationYingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma et al.KDD 2025 · 2 citations
- Rec: Towards Large Recommender Models with ReasoningRunyang You, Yongqi Li, Xinyu Lin, Xin Zhang et al.NeurIPS 2025 · 3 citations
- MSR-Rec: Multi-Step Reasoning-Enhanced LLM for Sequential RecommendationTuo Wang, Meng Jian, Ge Shi, Lifang Wu et al.AAAI 2026
- LEARN: Knowledge Adaptation from Large Language Model to Recommendation for Practical Industrial ApplicationJian Jia, Yipei Wang, Yan Li, Honggang Chen et al.AAAI 2025 · 37 citations
