Long-Sequence Recommendation Models Need Decoupled Embeddings
Ningya Feng, Junwei Pan, Jialong Wu, Baixu Chen, Ximei Wang, Qian Li, Xian Hu, Jie Jiang, Mingsheng Long
摘要
Lifelong user behavior sequences are crucial for capturing user interests and predicting user responses in modern recommendation systems. A two-stage paradigm is typically adopted to handle these long sequences: a subset of relevant behaviors is first searched from the original long sequences via an attention mechanism in the first stage and then aggregated with the target item to construct a discriminative representation for prediction in the second stage. In this work, we identify and characterize, for the first time, a neglected deficiency in existing long-sequence recommendation models: a single set of embeddings struggles with learning both attention and representation, leading to interference between these two processes. Initial attempts to address this issue with some common methods (e.g., linear projections-a technique borrowed from language processing) proved ineffective, shedding light on the unique challenges of recommendation models. To overcome this, we propose the Decoupled Attention and Representation Embeddings (DARE) model, where two distinct embedding tables are initialized and learned separately to fully decouple attention and representation. Extensive experiments and analysis demonstrate that DARE provides more accurate searches of correlated behaviors and outperforms baselines with AUC gains up to 9‰ on public datasets and notable improvements on Tencent's advertising platform. Furthermore, decoupling embedding spaces allows us to reduce the attention embedding dimension and accelerate the search procedure by 50% without significant performance impact, enabling more efficient, high-performance online serving. Code in PyTorch for experiments, including model analysis, is available at https://github.com/thuml/DARE . ˚Equal contribution. Work was done while Ningya Feng and Baixu Chen were interns at Tencent. 1 In this paper, "attention" refers to attention scores-the softmax output that weights each behavior.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language ModelsYuhao Wang, Junwei Pan, Pengyue Jia, Wanyu Wang 等SIGIR 2025 · 被引用 8 次
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate PredictionZixuan Li, Binzong Geng, Jing Xiong, Yong He 等KDD 2026 · 被引用 3 次
- Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR PredictionZhicheng Zhang, Zhaocheng Du, Jieming Zhu, Jiwei Tang 等AAAI 2026 · 被引用 2 次
- BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential RecommendationsMengyang Ma, Xiaopeng Li, Wanyu Wang, Zhaocheng Du 等WWW 2026 · 被引用 1 次
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone 等NeurIPS 2021 · 被引用 686 次
- Learning Optimal Tree Models under Beam SearchJingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu 等ICML 2020 · 被引用 72 次
相关 Paper
- Unleashing the Potential of Two-Tower Models: Diffusion-Based Cross-Interaction for Large-Scale MatchingYihan Wang, Fei Xiong, Zhexin Han, Qi Song 等WWW 2025 · 被引用 6 次
- When Search Meets Recommendation: Learning Disentangled Search Representation for RecommendationZihua Si, Zhongxiang Sun, Xiao Zhang, Jun Xu 等SIGIR 2023 · 被引用 29 次
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang 等WWW 2022 · 被引用 14 次
- Hierarchical Tree Search-based User Lifelong Behavior Modeling on Large Language ModelYu Xia, Rui Zhong, Hao Gu, Wei Yang 等SIGIR 2025 · 被引用 5 次
- Dynamic Memory based Attention Network for Sequential RecommendationQiaoyu Tan, Jianwei Zhang, Ninghao Liu, Xiao Huang 等AAAI 2021 · 被引用 74 次
