Reinforced Latent Reasoning for LLM-based Recommendation
Yang Zhang, Wenxin Xu, Xiaoyan Zhao, Wenjie Wang, Fuli Feng, Xiangnan He, Tat-Seng Chua
摘要
Large Language Models (LLMs) have demonstrated impressive reasoning capabilities in complex problem-solving tasks, sparking growing interest in their application to preference reasoning in recommendation systems. Existing methods typically rely on fine-tuning with explicit chain-of-thought (CoT) data. However, these methods face significant practical limitations due to (1) the difficulty of obtaining high-quality CoT data in recommendation and (2) the high inference latency caused by generating CoT reasoning. In this work, we explore an alternative approach that shifts from explicit CoT reasoning to compact, information-dense latent reasoning. This approach eliminates the need for explicit CoT generation and improves inference efficiency, as few latent tokens can effectively capture the entire reasoning process. Building on this idea, we propose Reinforced Latent Reasoning for Recommendation (LatentR), a novel end-to-end training framework that leverages reinforcement learning (RL) to optimize latent reasoning without relying on any CoT data. LatentR adopts a two-stage training strategy: first, supervised fine-tuning to initialize the latent reasoning module, followed by pure RL training to encourage exploration through a rule-based reward design. Our RL implementation is based on a modified GRPO algorithm, which reduces computational overhead during training and introduces continuous reward signals for more efficient learning. Extensive experiments demonstrate that LatentR enables effective latent reasoning without any direct supervision of the reasoning process, significantly improving performance when integrated with different LLM-based recommendation methods. Our codes are available at https://github.com/xuwenxinedu/R3 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Reasoning over Semantic IDs Enhances Generative RecommendationYingzhi He, Yan Sun, Junfei Tan, Yuxin Chen 等KDD 2026 · 被引用 15 次
- IGD: Token Decisiveness Modeling via Information Gain in LLMs for Personalized RecommendationZijie Lin, Yang Zhang, Xiaoyan Zhao, Fengbin Zhu 等NeurIPS 2025 · 被引用 12 次
- ThinkRec: Thinking-based recommendation via LLMQihang Yu, Kairui Fu, Zheqi Lv, Shengyu Zhang 等WWW 2026 · 被引用 10 次
- Show, Don't Tell: Morphing Latent Reasoning into Image GenerationHarold Haodong Chen, Xinxiang Yin, Wenjie Shu, Hongfei (Faye) Zhang 等ICML 2026 · 被引用 7 次
- Intuition-Guided Latent Reasoning for LLM-Based RecommendationChang Liu, Yimeng Bai, Xiaoyan Zhao, Yang Zhang 等KDD 2026 · 被引用 2 次
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- Think before you speak: Training Language Models With Pause TokensSachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon 等ICLR 2024 · 被引用 240 次
- AgentCF: Collaborative Learning with Autonomous Language Agents for Recommender SystemsJunjie Zhang, Yupeng Hou, Ruobing Xie, Wenqi Sun 等WWW 2024 · 被引用 164 次
- Large Language Models for Intent-Driven Session RecommendationsZhu Sun, Hongyang Liu, Xinghua Qu, Kaidong Feng 等SIGIR 2024 · 被引用 35 次
相关 Paper
- Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based RecommendationZhihao Ding, Jinming Li, Shuai Mu, Jieming ShiICLR 2026
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning ChainsWenhui Tan, Jiaze Li, Jianzhong Ju, Zhenbo Luo 等NeurIPS 2025 · 被引用 103 次
- CoT4Rec: Revealing User Preferences Through Chain of Thought for Recommender SystemsWeiqi Yue, Yuyu Yin, Xin Zhang, Binbin Shi 等AAAI 2025 · 被引用 10 次
- Learning to Reason over Continuous Tokens with Reinforcement LearningYiran Zhao, Yuhui Xu, Doyen Sahoo, Caiming Xiong 等ICLR 2026 · 被引用 1 次
- Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationJieyong Kim, Hyunseo Kim, Hyunjin Cho, SeongKu Kang 等SIGIR 2025 · 被引用 13 次
