Long-Sequence Recommendation Models Need Decoupled Embeddings
Ningya Feng, Junwei Pan, Jialong Wu, Baixu Chen, Ximei Wang, Qian Li, Xian Hu, Jie Jiang, Mingsheng Long
Abstract
Lifelong user behavior sequences are crucial for capturing user interests and predicting user responses in modern recommendation systems. A two-stage paradigm is typically adopted to handle these long sequences: a subset of relevant behaviors is first searched from the original long sequences via an attention mechanism in the first stage and then aggregated with the target item to construct a discriminative representation for prediction in the second stage. In this work, we identify and characterize, for the first time, a neglected deficiency in existing long-sequence recommendation models: a single set of embeddings struggles with learning both attention and representation, leading to interference between these two processes. Initial attempts to address this issue with some common methods (e.g., linear projections-a technique borrowed from language processing) proved ineffective, shedding light on the unique challenges of recommendation models. To overcome this, we propose the Decoupled Attention and Representation Embeddings (DARE) model, where two distinct embedding tables are initialized and learned separately to fully decouple attention and representation. Extensive experiments and analysis demonstrate that DARE provides more accurate searches of correlated behaviors and outperforms baselines with AUC gains up to 9‰ on public datasets and notable improvements on Tencent's advertising platform. Furthermore, decoupling embedding spaces allows us to reduce the attention embedding dimension and accelerate the search procedure by 50% without significant performance impact, enabling more efficient, high-performance online serving. Code in PyTorch for experiments, including model analysis, is available at https://github.com/thuml/DARE . ˚Equal contribution. Work was done while Ningya Feng and Baixu Chen were interns at Tencent. 1 In this paper, "attention" refers to attention scores-the softmax output that weights each behavior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Pre-train, Align, and Disentangle: Empowering Sequential Recommendation with Large Language ModelsYuhao Wang, Junwei Pan, Pengyue Jia, Wanyu Wang et al.SIGIR 2025 · 8 citations
- CTR-Sink: Attention Sink for Language Models in Click-Through Rate PredictionZixuan Li, Binzong Geng, Jing Xiong, Yong He et al.KDD 2026 · 3 citations
- Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR PredictionZhicheng Zhang, Zhaocheng Du, Jieming Zhu, Jiwei Tang et al.AAAI 2026 · 2 citations
- BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential RecommendationsMengyang Ma, Xiaopeng Li, Wanyu Wang, Zhaocheng Du et al.WWW 2026 · 1 citation
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Conflict-Averse Gradient Descent for Multi-task learningBo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone et al.NeurIPS 2021 · 686 citations
- Learning Optimal Tree Models under Beam SearchJingwei Zhuo, Ziru Xu, Wei Dai, Han Zhu et al.ICML 2020 · 72 citations
Related papers
- Unleashing the Potential of Two-Tower Models: Diffusion-Based Cross-Interaction for Large-Scale MatchingYihan Wang, Fei Xiong, Zhexin Han, Qi Song et al.WWW 2025 · 6 citations
- When Search Meets Recommendation: Learning Disentangled Search Representation for RecommendationZihua Si, Zhongxiang Sun, Xiao Zhang, Jun Xu et al.SIGIR 2023 · 29 citations
- Sequential Recommendation with Decomposed Item Feature RoutingKun Lin, Zhenlei Wang, Shiqi Shen, Zhipeng Wang et al.WWW 2022 · 14 citations
- Hierarchical Tree Search-based User Lifelong Behavior Modeling on Large Language ModelYu Xia, Rui Zhong, Hao Gu, Wei Yang et al.SIGIR 2025 · 5 citations
- Dynamic Memory based Attention Network for Sequential RecommendationQiaoyu Tan, Jianwei Zhang, Ninghao Liu, Xiao Huang et al.AAAI 2021 · 74 citations
