Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models
Tyler A. Chang, Yifan Xu, Weijian Xu, Zhuowen Tu
摘要
In this paper, we detail the relationship between convolutions and self-attention in natural language tasks. We show that relative position embeddings in self-attention layers are equivalent to recently-proposed dynamic lightweight convolutions, and we consider multiple new ways of integrating convolutions into Transformer self-attention. Specifically, we propose composite attention, which unites previous relative position embedding methods under a convolutional framework. We conduct experiments by training BERT with composite attention, finding that convolutions consistently improve performance on multiple downstream tasks, replacing absolute position embeddings. To inform future work, we present results comparing lightweight convolutions, dynamic convolutions, and depthwiseseparable convolutions in language model pretraining, considering multiple injection points for convolutions in self-attention layers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DKPLM: Decomposable Knowledge-Enhanced Pre-trained Language Model for Natural Language UnderstandingTaolin Zhang, Chengyu Wang, Nan Hu, Minghui Qiu 等AAAI 2022 · 被引用 36 次
- Monotonic Location Attention for Length GeneralizationJishnu Ray Chowdhury, Cornelia CarageaICML 2023 · 被引用 11 次
- Integral Transformer: Denoising Attention, Not Too Much Not Too LittleIvan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing ChenEMNLP 2025
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- On the Relationship between Self-Attention and Convolutional LayersJean-Baptiste Cordonnier, Andreas Loukas, Martin JaggiICLR 2020 · 被引用 629 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
相关 Paper
- ConvBERT: Improving BERT with Span-based Dynamic ConvolutionZihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen 等NeurIPS 2020 · 被引用 220 次
- A Simple and Effective Positional Encoding for TransformersPu-Chin Chen, Henry Tsai, Srinadh Bhojanapalli, Hyung Won Chung 等EMNLP 2021 · 被引用 51 次
- LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language ModelsHaoyu Wang, Ruirui Li, Haoming Jiang, Zhengyang Wang 等KDD 2023 · 被引用 5 次
- Rethinking Positional Encoding in Language Pre-trainingGuolin Ke, Di He, Tie-Yan LiuICLR 2021 · 被引用 358 次
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
