Are Pretrained Convolutions Better than Pretrained Transformers?
Yi Tay, Mostafa Dehghani, Jai Prakash Gupta, Vamsi Aribandi, Dara Bahri, Zhen Qin, Donald Metzler
Abstract
In the era of pre-trained language models, Transformers are the de facto choice of model architectures. While recent research has shown promise in entirely convolutional, or CNN, architectures, they have not been explored using the pre-train-fine-tune paradigm. In the context of language models, are convolutional models competitive to Transformers when pre-trained? This paper investigates this research question and presents several interesting findings. Across an extensive set of experiments on 8 datasets/tasks, we find that CNN-based pre-trained models are competitive and outperform their Transformer counterpart in certain scenarios, albeit with caveats. Overall, the findings outlined in this paper suggest that conflating pre-training and architectural advances is misguided and that both advances should be considered independently. We believe our research paves the way for a healthy amount of optimism in alternative architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Charformer: Fast Character Transformers via Gradient-based Subword TokenizationYi Tay, Vinh Q. Tran, Sebastian Ruder, Jai Prakash Gupta et al.ICLR 2022 · 198 citations
- On the Connection between Local Attention and Dynamic Depth-wise ConvolutionQi Han, Zejia Fan, Qi Dai, Lei Sun et al.ICLR 2022 · 144 citations
- UL2: Unifying Language Learning ParadigmsYi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia et al.ICLR 2023 · 97 citations
- Scale Efficiently: Insights from Pretraining and Finetuning TransformersYi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus et al.ICLR 2022 · 67 citations
- FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor CoresDaniel Y. Fu, Hermann Kumbong, Eric Nguyen, Christopher RéICLR 2024 · 41 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- Accelerating Large-Scale Inference with Anisotropic Vector QuantizationRuiqi Guo, Philip Sun, Erik Lindgren, Quan Geng et al.ICML 2020 · 539 citations
Related papers
- BERTAC: Enhancing Transformer-based Language Models with Adversarially Pretrained Convolutional Neural NetworksJong-Hoon Oh, Ryu Iida, Julien Kloetzer, Kentaro TorisawaACL 2021
- Training Object Detectors from Scratch: An Empirical Study in the Era of Vision TransformerWeixiang Hong, Jiangwei Lao, Wang Ren, Jian Wang et al.CVPR 2022 · 14 citations
- Are Transformers more robust than CNNs?Yutong Bai, Jieru Mei, Alan L. Yuille, Cihang XieNeurIPS 2021 · 365 citations
- Condenser: a Pre-training Architecture for Dense RetrievalLuyu Gao, Jamie CallanEMNLP 2021
- Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning AbilityIvan Lee, Nan Jiang, Taylor Berg-KirkpatrickICLR 2024 · 16 citations
