Implicit Kernel Attention
Kyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul Moon
Abstract
Attention computes the dependency between representations, and it encourages the model to focus on the important selective features. Attention-based models, such as Transformer and graph attention network (GAT), are widely utilized for sequential data and graph-structured data. This paper suggests a new interpretation and generalized structure of the attention in Transformer and GAT. For the attention in Transformer and GAT, we derive that the attention is a product of two parts: 1) the RBF kernel to measure the similarity of two instances and 2) the exponential of L 2 norm to compute the importance of individual instances. From this decomposition, we generalize the attention in three ways. First, we propose implicit kernel attention with an implicit kernel function instead of manual kernel selection. Second, we generalize L 2 norm as the L p norm. Third, we extend our attention to structured multi-head attention. Our generalized attention shows better performance on classification, translation, and regression tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfd2cd0b-d1ff-4d6e-8c78-2f217792a3a8Cited by top-tier papers4
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 516 citations
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li et al.ICLR 2021 · 171 citations
- Uniform Memory Retrieval with Larger Capacity for Modern Hopfield ModelsDennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, Han LiuICML 2024 · 44 citations
- Perceiving Longer Sequences With Bi-Directional Cross-Attention TransformersMarkus Hiller, Krista A. Ehinger, Tom DrummondNeurIPS 2024 · 23 citations
Builds on3
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li et al.ICLR 2020 · 134 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
Related papers
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen et al.NeurIPS 2022 · 59 citations
- A Primal-Dual Framework for Transformers and Neural NetworksTan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi et al.ICLR 2023 · 2 citations
- DHAKR: Learning Deep Hierarchical Attention-Based Kernelized Representations for Graph ClassificationFeifei Qian, Lu Bai, Lixin Cui, Ming Li et al.AAAI 2025 · 4 citations
- Sequential Recommendation with Relation-Aware Kernelized Self-AttentionMingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim et al.AAAI 2020 · 31 citations
- How to Find Your Friendly Neighborhood: Graph Attention Design with Self-SupervisionDongkwan Kim, Alice OhICLR 2021 · 309 citations
