CollectiveKV: Decoupling and Sharing Collaborative Information in Sequential Recommendation
Jingyu Li, Zhaocheng Du, Qianhui Zhu, Kaiyuan Li, Zhicheng Zhang, Song-Li Wu, Chaolang Li, Pengwen Dai
摘要
Sequential recommendation models are widely used in applications, yet they face stringent latency requirements. Mainstream models leverage the Transformer attention mechanism to improve performance, but its computational complexity grows with the sequence length, leading to a latency challenge for long sequences. Consequently, KV cache technology has recently been explored in sequential recommendation systems to reduce inference latency. However, KV cache introduces substantial storage overhead in sequential recommendation systems, which often have a large user base with potentially very long user history sequences. In this work, we observe that KV sequences across different users exhibit significant similarities, indicating the existence of collaborative signals in KV. Furthermore, we analyze the KV using singular value decomposition (SVD) and find that the information in KV can be divided into two parts: the majority of the information is shareable across users, while a small portion is user-specific. Motivated by this, we propose CollectiveKV, a cross-user KV sharing mechanism. It captures the information shared across users through a learnable global KV pool. During inference, each user retrieves high-dimensional shared KV from the pool and concatenates them with low-dimensional user-specific KV to obtain the final KV. Experiments on five sequential recommendation models and three datasets show that our method can compress the KV cache to only 0.8% of its original size, while maintaining or even enhancing model performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- QUEST: Query-Aware Sparsity for Efficient Long-Context LLM InferenceJiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao 等ICML 2024 · 被引用 316 次
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache ManagementWonbeom Lee, Jungi Lee, Junghwan Seo, Jaewoong SimOSDI 2024 · 被引用 248 次
- Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative RecommendationsJiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang 等ICML 2024 · 被引用 200 次
- Loki: Low-rank Keys for Efficient Sparse AttentionPrajwal Singhania, Siddharth Singh, Shwai He, Soheil Feizi 等NeurIPS 2024 · 被引用 94 次
- Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR PredictionZhicheng Zhang, Zhaocheng Du, Jieming Zhu, Jiwei Tang 等AAAI 2026 · 被引用 2 次
相关 Paper
- EARN: Efficient Inference Acceleration for LLM-based Generative Recommendation by Register TokensChaoqun Yang, Xinyu Lin, Wenjie Wang, Yongqi Li 等KDD 2025 · 被引用 1 次
- Breaking the Bottleneck: User-Specific Optimization and Real-Time Inference Integration for Sequential RecommendationWenjia Xie, Hao Wang, Minghao Fang, Ruize Yu 等KDD 2025
- Sparse Attention Across Multiple-Context KV CacheZiyi Cao, Qingyi Si, Jingbin Zhang, Bingquan LiuAAAI 2026 · 被引用 3 次
- Collaborative Memory Augmentation for Generative RecommendationEnze Liu, Zhen Tian, Wayne Xin ZhaoKDD 2026 · 被引用 1 次
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector ExtractionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang 等ICML 2026 · 被引用 3 次
