Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation
Chao Chen, Haoyu Geng, Nianzu Yang, Junchi Yan, Daiyue Xue, Jianping Yu, Xiaokang Yang
Abstract
User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existing user behavior modeling solutions, attention networks are widely adopted for its effectiveness and relative simplicity. Despite being extensively studied, existing attentions still suffer from two limitations: i) conventional attentions mainly take into account the spatial correlation between user behaviors, regardless the distance between those behaviors in the continuous time space; and ii) these attentions mostly provide a dense and undistinguished distribution over all past behaviors then attentively encode them into the output latent representations. This is however not suitable in practical scenarios where a user's future actions are relevant to a small subset of her/his historical behaviors. In this paper, we propose a novel attention network, named self-modulating attention, that models the complex and non-linearly evolving dynamic user preferences. We empirically demonstrate the effectiveness of our method on top-N sequential recommendation tasks, and the results on three large-scale real-world datasets show that our model can achieve state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ContiFormer: Continuous-Time Transformer for Irregular Time Series ModelingYuqi Chen, Kan Ren, Yansen Wang, Yuchen Fang et al.NeurIPS 2023 · 131 citations
- Recommender Transformers with Behavior PathwaysZhiyu Yao, Xinyang Chen, Sinan Wang, Qinyan Dai et al.WWW 2024 · 9 citations
Builds on6
- Transformer Hawkes ProcessSimiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao et al.ICML 2020 · 382 citations
- Multi-Time Attention Networks for Irregularly Sampled Time SeriesSatya Narayan Shukla, Benjamin M. MarlinICLR 2021 · 301 citations
- Self-Attentive Hawkes ProcessQiang Zhang, Aldo Lipani, Ömer Kirnap, Emine YilmazICML 2020 · 254 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
- Kalman Filtering Attention for User Behavior Modeling in CTR PredictionHu Liu, Jing Lu, Xiwei Zhao, Sulong Xu et al.NeurIPS 2020 · 29 citations
Related papers
- Déjà vu: A Contextualized Temporal Attention Mechanism for Sequential RecommendationJibang Wu, Renqin Cai, Hongning WangWWW 2020 · 66 citations
- Dynamic Memory based Attention Network for Sequential RecommendationQiaoyu Tan, Jianwei Zhang, Ninghao Liu, Xiao Huang et al.AAAI 2021 · 74 citations
- Why Generate When You Can Transform? Unleashing Generative Attention for Dynamic RecommendationYuli Liu, Wenjun Kong, Weizhi Ma, Cheng LuoACM MM 2025
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu et al.SIGIR 2023 · 142 citations
- Modeling Stage-wise Evolution of User Interests for News RecommendationZhiyong Cheng, Yike Jin, Zhijie Zhang, Huilin Chen et al.WWW 2026
