FourierFormer: Transformer Meets Generalized Fourier Integral Theorem
Tan Nguyen, Minh Pham, Tam Nguyen, Khai Nguyen, Stanley J. Osher, Nhat Ho
Abstract
Multi-head attention empowers the recent success of transformers, the state-of-theart models that have achieved remarkable success in sequence modeling and beyond. These attention mechanisms compute the pairwise dot products between the queries and keys, which results from the use of unnormalized Gaussian kernels with the assumption that the queries follow a mixture of Gaussian distribution. There is no guarantee that this assumption is valid in practice. In response, we first interpret attention in transformers as a nonparametric kernel regression. We then propose the FourierFormer, a new class of transformers in which the dot-product kernels are replaced by the novel generalized Fourier integral kernels. Different from the dot-product kernels, where we need to choose a good covariance matrix to capture the dependency of the features of data, the generalized Fourier integral kernels can automatically capture such dependency and remove the need to tune the covariance matrix. We theoretically prove that our proposed Fourier integral kernels can efficiently approximate any key and query distributions. Compared to the conventional transformers with dot-product attention, FourierFormers attain better accuracy and reduce the redundancy between attention heads. We empirically corroborate the advantages of FourierFormers over the baseline transformers in a variety of practical applications including language modeling and image classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5154290a-d356-41db-a889-a597f7ea7829Cited by top-tier papers17
- Scalable Transformer for PDE Surrogate ModelingZijie Li, Dule Shu, Amir Barati FarimaniNeurIPS 2023 · 188 citations
- Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In ContextXiang Cheng, Yuxin Chen, Suvrit SraICML 2024 · 64 citations
- Visual Fourier Prompt TuningRunjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu et al.NeurIPS 2024 · 58 citations
- Primal-Attention: Self-attention through Asymmetric Kernel SVD in Primal RepresentationYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensNeurIPS 2023 · 42 citations
- Improving Transformer with an Admixture of Attention HeadsTan Nguyen, Tam Nguyen, Hai Do, Khai Nguyen et al.NeurIPS 2022 · 38 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen et al.ICML 2022 · 38 citations
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 18 citations
- Dense-Exponential Random Features: Sharp Positive Estimators of the Gaussian KernelValerii Likhosherstov, Krzysztof Marcin Choromanski, Kumar Avinava Dubey, Frederick Liu et al.NeurIPS 2023 · 5 citations
- KDEformer: Accelerating Transformers via Kernel Density EstimationAmir Zandieh, Insu Han, Majid Daliri, Amin KarbasiICML 2023 · 55 citations
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song et al.ICLR 2021 · 122 citations
