Implicit Kernel Attention
Kyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul Moon
摘要
Attention computes the dependency between representations, and it encourages the model to focus on the important selective features. Attention-based models, such as Transformer and graph attention network (GAT), are widely utilized for sequential data and graph-structured data. This paper suggests a new interpretation and generalized structure of the attention in Transformer and GAT. For the attention in Transformer and GAT, we derive that the attention is a product of two parts: 1) the RBF kernel to measure the similarity of two instances and 2) the exponential of L 2 norm to compute the importance of individual instances. From this decomposition, we generalize the attention in three ways. First, we propose implicit kernel attention with an implicit kernel function instead of manual kernel selection. Second, we generalize L 2 norm as the L p norm. Third, we extend our attention to structured multi-head attention. Our generalized attention shows better performance on classification, translation, and regression tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 被引用 516 次
- Is Attention Better Than Matrix Decomposition?Zhengyang Geng, Meng-Hao Guo, Hongxu Chen, Xia Li 等ICLR 2021 · 被引用 171 次
- Uniform Memory Retrieval with Larger Capacity for Modern Hopfield ModelsDennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, Han LiuICML 2024 · 被引用 44 次
- Perceiving Longer Sequences With Bi-Directional Cross-Attention TransformersMarkus Hiller, Krista A. Ehinger, Tom DrummondNeurIPS 2024 · 被引用 23 次
它引用的顶会 Paper3
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Encoding word order in complex embeddingsBenyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li 等ICLR 2020 · 被引用 134 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
相关 Paper
- FourierFormer: Transformer Meets Generalized Fourier Integral TheoremTan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen 等NeurIPS 2022 · 被引用 59 次
- A Primal-Dual Framework for Transformers and Neural NetworksTan Minh Nguyen, Tam Minh Nguyen, Nhat Ho, Andrea L. Bertozzi 等ICLR 2023 · 被引用 2 次
- DHAKR: Learning Deep Hierarchical Attention-Based Kernelized Representations for Graph ClassificationFeifei Qian, Lu Bai, Lixin Cui, Ming Li 等AAAI 2025 · 被引用 4 次
- Sequential Recommendation with Relation-Aware Kernelized Self-AttentionMingi Ji, Weonyoung Joo, Kyungwoo Song, Yoon-Yeong Kim 等AAAI 2020 · 被引用 31 次
- How to Find Your Friendly Neighborhood: Graph Attention Design with Self-SupervisionDongkwan Kim, Alice OhICLR 2021 · 被引用 309 次
