An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention
Yehjin Shin, Jeongwhan Choi, Hyowon Wi, Noseong Park
Abstract
Sequential recommendation (SR) models based on Transformers have achieved remarkable successes. The self-attention mechanism of Transformers for computer vision and natural language processing suffers from the oversmoothing problem, i.e., hidden representations becoming similar to tokens. In the SR domain, we, for the first time, show that the same problem occurs. We present pioneering investigations that reveal the low-pass filtering nature of self-attention in the SR, which causes oversmoothing. To this end, we propose a novel method called Beyond Self-Attention for Sequential Recommendation (BSARec), which leverages the Fourier transform to i) inject an inductive bias by considering fine-grained sequential patterns and ii) integrate low and high-frequency information to mitigate oversmoothing. Our discovery shows significant advancements in the SR domain and is expected to bridge the gap for existing Transformer-based SR models. We test our proposed approach through extensive experiments on 6 benchmark datasets. The experimental results demonstrate that our model outperforms 7 baseline methods in terms of recommendation performance. Our code is available at https://github.com/yehjin-shin/BSARec.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5f32d9f-99ea-48e2-b610-e1d57b029ea3Cited by top-tier papers21
- Graph Convolutions Enrich the Self-Attention in Transformers!Jeongwhan Choi, Hyowon Wi, Jayoung Kim, Yehjin Shin et al.NeurIPS 2024 · 24 citations
- Structured Spectral Reasoning for Frequency-Adaptive Multimodal RecommendationWei Yang, Rui Zhong, Yiqun Chen, Chi Lu et al.NeurIPS 2025 · 10 citations
- Enhancing Long-and Short-Term Representations for Next POI Recommendations via Frequency and Hierarchical Contrastive LearningJiajie Chen, Yu Sang, Peng-Fei Zhang, Jiaan Wang et al.AAAI 2025 · 9 citations
- CSRec: Rethinking Sequential Recommendation from A Causal PerspectiveXiaoyu Liu, Jiaxin Yuan, Yuhang Zhou, Jingling Li et al.SIGIR 2025 · 7 citations
- FIM: Frequency-Aware Multi-View Interest Modeling for Local-Life Service RecommendationGuoquan Wang, Qiang Luo, Weisong Hu, Pengfei Yao et al.SIGIR 2025 · 7 citations
Builds on10
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Attention is not all you need: pure attention loses rank doubly exponentially with depthYihe Dong, Jean-Baptiste Cordonnier, Andreas LoukasICML 2021 · 522 citations
- On Sampled Metrics for Item RecommendationWalid Krichene, Steffen RendleKDD 2020 · 459 citations
- Filter-enhanced MLP is All You Need for Sequential RecommendationKun Zhou, Hui Yu, Wayne Xin Zhao, Ji-Rong WenWWW 2022 · 411 citations
- Anti-Oversmoothing in Deep Vision Transformers via the Fourier Domain Analysis: From Theory to PracticePeihao Wang, Wenqing Zheng, Tianlong Chen, Zhangyang WangICLR 2022 · 212 citations
Related papers
- Frequency Enhanced Hybrid Attention Network for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Jianfeng Qu et al.SIGIR 2023 · 142 citations
- Wavelet Enhanced Adaptive Frequency Filter for Sequential RecommendationHuayang Xu, Huanhuan Yuan, Guanfeng Liu, Junhua Fang et al.AAAI 2026 · 1 citation
- Why Generate When You Can Transform? Unleashing Generative Attention for Dynamic RecommendationYuli Liu, Wenjun Kong, Weizhi Ma, Cheng LuoACM MM 2025
- Contrastive Enhanced Slide Filter Mixer for Sequential RecommendationXinyu Du, Huanhuan Yuan, Pengpeng Zhao, Junhua Fang et al.ICDE 2023 · 18 citations
- Recommender Transformers with Behavior PathwaysZhiyu Yao, Xinyang Chen, Sinan Wang, Qinyan Dai et al.WWW 2024 · 9 citations
