Token-Efficient Long-Term Interest Sketching and Internalized Reasoning for LLM-based Recommendation
Zhihao Ding, Jinming Li, Shuai Mu, Jieming Shi
Abstract
Large language models (LLMs) can solve complex real-world tasks when prompted to generate chain-of-thought (CoT) reasoning, motivating their use for preference reasoning in recommender systems. However, applying LLM reasoning on recommendation faces two practical challenges. First, LLMs struggle to reason over long, noisy user histories that often span hundreds of items while truncation discards signals needed to capture long-term interests. Second, in decoder-only architectures, CoT requires generating rationale tokens autoregressively, leading to prohibitive inference latency for real-world deployment. To address the challenges, we propose SIREN, a framework that enables effective LLM-based rating prediction via long-term interest sketching and internalized reasoning. First, instead of prompting raw histories, we build a compact, token-bounded interest sketch that preserves persistent preferences and suppresses noise. Specifically, we encode and cluster item descriptions to discover semantic topics, then compress each user’s history into a short list of liked and disliked topics, facilitating LLM reasoning. Second, we develop an internalized reasoning strategy for efficient inference. We adopt a two-stage training paradigm: (i) train the LLM to reason explicitly for rating prediction with rule-based reinforcement learning, since ground-truth CoTs are unavailable in recommendation; and (ii) learn to internalize CoT into model parameters through hidden alignment. At inference, the LLM directly generates the rating with near-CoT quality. Extensive experiments show that SIREN reduces average input tokens by compared to raw-history prompting, outperforms existing methods while delivering over lower inference latency than CoT-based LLM recommenders. Code and data are available at https://github.com/TommyDzh/SIREN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
Related papers
- Reinforced Latent Reasoning for LLM-based RecommendationYang Zhang, Wenxin Xu, Xiaoyan Zhao, Wenjie Wang et al.ICLR 2026 · 67 citations
- CoT4Rec: Revealing User Preferences Through Chain of Thought for Recommender SystemsWeiqi Yue, Yuyu Yin, Xin Zhang, Binbin Shi et al.AAAI 2025 · 10 citations
- Think, But Don't Tell: Implicit Reasoning for LLM-based Sequential Recommendation via Multi-Teacher DistillationWeihai Lu, Xiaoxi Cui, Chenke YinSIGIR 2026
- Intuition-Guided Latent Reasoning for LLM-Based RecommendationChang Liu, Yimeng Bai, Xiaoyan Zhao, Yang Zhang et al.KDD 2026 · 2 citations
- ThinkRec: Thinking-based recommendation via LLMQihang Yu, Kairui Fu, Zheqi Lv, Shengyu Zhang et al.WWW 2026 · 10 citations
