CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
Zixuan Li, Binzong Geng, Jing Xiong, Yong He, Yuxuan Hu, Jian Chen, Dingwei Chen, Xiyu Chang, Ngai Wong, Liang Zhang, Linjian Mo, Chengming Li
Abstract
Click-Through Rate (CTR) prediction, a core task in recommendation systems, estimates user click likelihood using historical behavioral data. Modeling user behavior sequences as text to leverage Language Models (LMs) for this task has gained traction, owing to LMs' strong semantic understanding and contextual modeling capabilities. However, a critical structural gap exists: user behavior sequences consist of discrete actions connected by semantically empty separators, differing fundamentally from the coherent natural language in LM pre-training. This mismatch causes semantic fragmentation, where the lack of syntactic coherence causes LM attention to scatters across irrelevant tokens instead of focusing on meaningful behavior boundaries and inter-behavior relationships, degrading prediction performance. To address this, we propose CTR-Sink, a novel framework introducing behavior-level attention sinks tailored for recommendation scenarios. Inspired by attention sink theory, it constructs attention focus sinks and dynamically regulates attention aggregation via external information. Specifically, we insert sink tokens between consecutive behaviors, incorporating recommendation-specific signals such as temporal distance to serve as stable attention sinks. To enhance generality, we design * This work was done during the internship at Ant Group. † The corresponding authors. a two-stage training strategy that explicitly guides LM attention toward sink tokens and an attention sink mechanism that amplifies inter-sink dependencies to better capture behavioral correlations. Experiments on one industrial dataset and two open-source datasets (MovieLens, Kuairec), alongside visualization results, validate the method's effectiveness across scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain et al.WWW 2021 · 793 citations
- ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationJianghao Lin, Rong Shan, Chenxu Zhu, Kounianhua Du et al.WWW 2024 · 151 citations
- Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention CalibrationZhongzhi Yu, Zheng Wang, Yonggan Fu, Huihong Shi et al.ICML 2024 · 63 citations
Related papers
- ClickPrompt: CTR Models are Strong Prompt Generators for Adapting Language Models to CTR PredictionJianghao Lin, Bo Chen, Hangyu Wang, Yunjia Xi et al.WWW 2024 · 58 citations
- Topic Guided Multi-faceted Semantic Disentanglement for CTR predictionFengxin Li, Zhiqian Yin, Hongyan Liu, Jingcai Guo et al.ACM MM 2025
- Field Matters: A Lightweight LLM-enhanced Method for CTR PredictionYu Cui, Feng Liu, Jiawei Chen, Xingyu Lou et al.WWW 2026 · 5 citations
- GenCI: Generative Modeling of User Interest Shift via Cohort-based Intent Learning for CTR PredictionKesha Ou, Zhen Tian, Wayne Xin Zhao, Hongyu Lu et al.WWW 2026
- Length-Adaptive Interest Network for Balancing Long and Short Sequence Modeling in CTR PredictionZhicheng Zhang, Zhaocheng Du, Jieming Zhu, Jiwei Tang et al.AAAI 2026 · 2 citations
