Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based Recommendation
Zhihao Ding, Jinming Li, Shuai Mu, Jieming Shi
摘要
Large language models (LLMs) can solve complex real-world tasks when prompted to generate chain-of-thought (CoT) reasoning, motivating their use for preference reasoning in recommender systems. However, applying LLM reasoning on recommendation faces two practical challenges. First, LLMs struggle to reason over long, noisy user histories that often span hundreds of items while truncation discards signals needed to capture long-term interests. Second, in decoder-only architectures, CoT requires generating rationale tokens autoregressively, leading to prohibitive inference latency for real-world deployment. To address the challenges, we propose SIREN, a framework that enables effective LLM-based rating prediction via long-term interest sketching and internalized reasoning. First, instead of prompting raw histories, we build a compact, token-bounded interest sketch that preserves persistent preferences and suppresses noise. Specifically, we encode and cluster item descriptions to discover semantic topics, then compress each user’s history into a short list of liked and disliked topics, facilitating LLM reasoning. Second, we develop an internalized reasoning strategy for efficient inference. We adopt a two-stage training paradigm: (i) train the LLM to reason explicitly for rating prediction with rule-based reinforcement learning, since ground-truth CoTs are unavailable in recommendation; and (ii) learn to internalize CoT into model parameters through hidden alignment. At inference, the LLM directly generates the rating with near-CoT quality. Extensive experiments show that SIREN reduces average input tokens by compared to raw-history prompting, outperforms existing methods while delivering over lower inference latency than CoT-based LLM recommenders. Code and data are available at https://github.com/TommyDzh/SIREN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
相关 Paper
- Reinforced Latent Reasoning for LLM-based RecommendationYang Zhang, Wenxin Xu, Xiaoyan Zhao, Wenjie Wang 等ICLR 2026 · 被引用 67 次
- CoT4Rec: Revealing User Preferences Through Chain of Thought for Recommender SystemsWeiqi Yue, Yuyu Yin, Xin Zhang, Binbin Shi 等AAAI 2025 · 被引用 10 次
- Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher DistillationWeihai Lu, Xiaoxi Cui, Chenke YinSIGIR 2026
- Intuition-Guided Latent Reasoning for LLM-Based RecommendationChang Liu, Yimeng Bai, Xiaoyan Zhao, Yang Zhang 等KDD 2026 · 被引用 2 次
- ThinkRec: Thinking-based recommendation via LLMQihang Yu, Kairui Fu, Zheqi Lv, Shengyu Zhang 等WWW 2026 · 被引用 10 次
