Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
Yaochen Zhu, Harald Steck, Dawen Liang, Yinhan He, Vito Claudio Ostuni, Jundong Li, Nathan Kallus
摘要
Large language models (LLMs) are reshaping the recommender system paradigm by enabling users to express preferences and receive recommendations through conversations. Yet, aligning LLMs to the recommendation task remains challenging: pretrained LLMs often generate out-of-catalog items, violate required output formats, and their ranking quality degrades sharply toward the end of the generated list. To this end, we propose ConvRec-R1, a two-stage framework for end-to-end training of LLM-based conversational recommender systems. In Stage 1, we construct a behavioral-cloning dataset with a Remap-Reflect-Adjust pipeline, which produces high-quality, catalog-grounded demonstrations from powerful blackbox LLMs to warm-start the RL training. In Stage 2, we propose Rank-GRPO, a principled extension of group relative policy optimization (GRPO) (Shao et al., 2024) tailored to tasks with rank-style outputs. Rank-GRPO treats each rank in the recommendation list as the unit instead of token (too fine-grained) or sequence (too coarse), redefining rewards to remove non-causal credit assignment and introducing a rank-level importance ratio based on the geometric mean of rankwise token probabilities to stabilize policy updates. Experiments on the public REDDIT-V2 dataset show that ConvRec-R1 converges faster and achieves higher Recall and NDCG than GRPO-style baselines. Code and datasets are released at https://github.com/yaochenzhu/Rank-GRPO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender SystemsHaochang Hao, Yifan Xu, Xinzhuo Li, Yingqiang Ge 等KDD 2026 · 被引用 1 次
- Reforming the Mechanism: Editing Reasoning Patterns in LLMs with Circuit ReshapingZhenyu Lei, Qiong Wu, JIANXIONG DONG, Yinhan He 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- Think before Recommendation: Autonomous Reasoning-enhanced RecommenderXiaoyu Kong, Junguang Jiang, Bin Liu, Ziru Xu 等NeurIPS 2025 · 被引用 17 次
- From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement LearningWenzhe Niu, Wei He, Zongxia Xie, Jinpeng Ou 等ICML 2026 · 被引用 1 次
- ToolRL: Reward is All Tool Learning NeedsCheng Qian, Emre Can Acikgoz, Qi He, Hongru Wang 等NeurIPS 2025 · 被引用 387 次
- Collaborative Retrieval for Large Language Model-based Conversational Recommender SystemsYaochen Zhu, Chao Wan, Harald Steck, Dawen Liang 等WWW 2025 · 被引用 15 次
- Refining Text Generation for Realistic Conversational Recommendation via Direct Preference OptimizationManato Tajiri, Michimasa InabaEMNLP 2025
