Linear-Time Self Attention with Codeword Histogram for Efficient Recommendation
Yongji Wu, Defu Lian, Neil Zhenqiang Gong, Lu Yin, Mingyang Yin, Jingren Zhou, Hongxia Yang
Abstract
Self-attention has become increasingly popular in a variety of sequence modeling tasks from natural language processing to recommendation, due to its effectiveness. However, self-attention suffers from quadratic computational and memory complexities, prohibiting its applications on long sequences. Existing approaches that address this issue mainly rely on a sparse attention context, either using a local window, or a permuted bucket obtained by localitysensitive hashing (LSH) or sorting, while crucial information may be lost. Inspired by the idea of vector quantization that uses cluster centroids to approximate items, we propose LISA (LInear-time Self Attention), which enjoys both the effectiveness of vanilla selfattention and the efficiency of sparse attention. LISA scales linearly with the sequence length, while enabling full contextual attention via computing differentiable histograms of codeword distributions. Meanwhile, unlike some efficient attention methods, our method poses no restriction on casual masking or sequence length. We evaluate our method on four real-world datasets for sequential recommendation. The results show that LISA outperforms the state-ofthe-art efficient attention methods in both performance and speed; and it is up to 57x faster and 78x more memory efficient than vanilla self-attention. CCS CONCEPTS • Information systems → Recommender systems; Users and interactive retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90d597ea-01e2-4d91-ac23-3a75e8656608Cited by top-tier papers4
- Learning Vector-Quantized Item Representation for Transferable Sequential RecommendersYupeng Hou, Zhankui He, Julian J. McAuley, Wayne Xin ZhaoWWW 2023 · 256 citations
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense EmbeddingsShitao Xiao, Zheng Liu, Weihao Han, Jianjin Zhang et al.SIGIR 2022 · 31 citations
- Cooperative Retriever and Ranker in Deep RecommendersXu Huang, Defu Lian, Jin Chen, Zheng Liu et al.WWW 2023 · 17 citations
- Balanced Co-Clustering of Users and Items for Embedding Table Compression in Recommender SystemsRunhao Jiang, Renchi Yang, Donghao WuSIGIR 2026
Builds on9
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- LayoutLM: Pre-training of Text and Layout for Document Image UnderstandingYiheng Xu, Minghao Li, Lei Cui, Shaohan Huang et al.KDD 2020 · 575 citations
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler et al.ICML 2020 · 391 citations
- Geography-Aware Sequential Location RecommendationDefu Lian, Yongji Wu, Yong Ge, Xing Xie et al.KDD 2020 · 244 citations
Related papers
- You Only Sample (Almost) Once: Linear Cost Self-Attention Via Bernoulli SamplingZhanpeng Zeng, Yunyang Xiong, Sathya N. Ravi, Shailesh Acharya et al.ICML 2021 · 22 citations
- Overcoming Long Context Limitations of State Space Models via Context Dependent Sparse AttentionZhihao Zhan, Jianan Zhao, Zhaocheng Zhu, Jian TangNeurIPS 2025 · 7 citations
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao et al.SIGIR 2023 · 86 citations
- Sparse Attention with Learning to HashZhiqing Sun, Yiming Yang, Shinjae YooICLR 2022 · 21 citations
- Iterative Sparse Attention for Long-sequence RecommendationGuanyu Lin, Jinwei Luo, Yinfeng Li, Chen Gao et al.AAAI 2025 · 2 citations
