Synthesizer: Rethinking Self-Attention for Transformer Models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, Che Zheng
Abstract
The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-based self-attention mechanism on the performance of Transformer models. Via extensive experiments, we find that (1) random alignment matrices surprisingly perform quite competitively and (2) learning attention weights from token-token (query-key) interactions is useful but not that important after all. To this end, we propose SYNTHESIZER, a model that learns synthetic attention weights without token-token interactions. In our experiments, we first show that simple Synthesizers achieve highly competitive performance when compared against vanilla Transformer models across a range of tasks, including machine translation, language modeling, text generation and GLUE/SuperGLUE benchmarks. When composed with dot product attention, we find that Synthesizers consistently outperform Transformers. Moreover, we conduct additional comparisons of Synthesizers against Dynamic Convolutions, showing that simple Random Synthesizer is not only 60% faster but also improves perplexity by a relative 3.5%. Finally, we show that simple factorized Synthesizers can outperform Linformers on encoding only tasks. introduction Transformer models (Vaswani et al., 2017) have demonstrated success across a wide range of tasks. This has resulted in Transformers largely displacing once popular auto-regressive and recurrent models in recent years. At the heart of Transformer models lies the query-key-value dot
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fabf139b-0a0c-43fb-a26d-0ce3ccb986e4Cited by top-tier papers46
- cosFormer: Rethinking Softmax In AttentionZhen Qin, Weixuan Sun, Hui Deng, Dongxu Li et al.ICLR 2022 · 303 citations
- Hierarchically Gated Recurrent Neural Network for Sequence ModelingZhen Qin, Songlin Yang, Yiran ZhongNeurIPS 2023 · 152 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- Container: Context Aggregation NetworksPeng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi et al.NeurIPS 2021 · 86 citations
Builds on4
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 629 citations
- Sparse Sinkhorn AttentionYi Tay, Dara Bahri, Liu Yang, Donald Metzler et al.ICML 2020 · 391 citations
Related papers
- Recurrent Attention for Neural Machine TranslationJiali Zeng, Shuangzhi Wu, Yongjing Yin, Yufan Jiang et al.EMNLP 2021
- Is the Attention Matrix Really the Key to Self-Attention in Multivariate Long-Term Time Series Forecasting?Xinyu Li, Kexi Chen, Jiajie Shen, Ying Zheng et al.ACL 2026
- Radar: Fast Long-Context Decoding for Any TransformerYongchang Hao, Mengyao Zhai, Hossein Hajimirsadeghi, Sepidehsadat Hosseini et al.ICLR 2025
- Hard-Coded Gaussian Attention for Neural Machine TranslationWeiqiu You, Simeng Sun, Mohit IyyerACL 2020 · 55 citations
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 28 citations
