What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?
Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, Colin Raffel
Abstract
Large pretrained Transformer language models have been shown to exhibit zeroshot generalization, i.e. they can perform a wide variety of tasks that they were not explicitly trained on. However, the architectures and pretraining objectives used across state-of-the-art models differ significantly, and there has been limited systematic comparison of these factors. In this work, we present a large-scale evaluation of modeling choices and their impact on zero-shot generalization. In particular, we focus on text-to-text models and experiment with three model architectures (causal/non-causal decoder-only and encoder-decoder), trained with two different pretraining objectives (autoregressive and masked language modeling), and evaluated with and without multitask prompted finetuning. We train models with over 5 billion parameters for more than 170 billion tokens, thereby increasing the likelihood that our conclusions will transfer to even larger scales. Our experiments show that causal decoder-only models trained on an autoregressive language modeling objective exhibit the strongest zero-shot generalization after purely unsupervised pretraining. However, models with non-causal visibility on their input trained with a masked language modeling objective followed by multitask finetuning perform the best among our experiments. We therefore consider the adaptation of pretrained models across architectures and objectives. We find that pretrained non-causal decoder models can be adapted into performant generative causal decoder models, using autoregressive language modeling as a downstream task. Furthermore, we find that pretrained causal decoder models can be efficiently adapted into non-causal decoder models, ultimately achieving competitive performance after multitask finetuning. Code and checkpoints are available at https://github.com/bigscience-workshop/architecture-objective . * Equal contribution. † Equal supervision. Individual contributions outlined in Appendix A
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers32
- AutoTimes: Autoregressive Time Series Forecasters via Large Language ModelsYong Liu, Guo Qin, Xiangdong Huang, Jianmin Wang et al.NeurIPS 2024 · 138 citations
- A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language ModelsHaoran Xu, Young Jin Kim, Amr Sharaf, Hany Hassan AwadallaICLR 2024 · 122 citations
- Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary SpaceMor Geva, Avi Caciularu, Kevin Ro Wang, Yoav GoldbergEMNLP 2022 · 92 citations
- One-Step Diffusion Distillation via Deep Equilibrium ModelsZhengyang Geng, Ashwini Pokle, J. Zico KolterNeurIPS 2023 · 83 citations
- VIMA: Robot Manipulation with Multimodal PromptsYunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang et al.ICML 2023 · 80 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- Examining Scaling and Transfer of Language Model Architectures for Machine TranslationBiao Zhang, Behrooz Ghorbani, Ankur Bapna, Yong Cheng et al.ICML 2022 · 30 citations
- A Universal Discriminator for Zero-Shot GeneralizationHaike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng et al.ACL 2023 · 6 citations
- Context-Aware Multimodal PretrainingKarsten Roth, Zeynep Akata, Dima Damen, Ivana Balazevic et al.CVPR 2025
- BERTGen: Multi-task Generation through BERTFaidon Mitzalis, Ozan Caglayan, Pranava Madhyastha, Lucia SpeciaACL 2021
- LaFTer: Label-Free Tuning of Zero-shot Classifier using Language and Unlabeled Image CollectionsMuhammad Jehanzeb Mirza, Leonid Karlinsky, Wei Lin, Horst Possegger et al.NeurIPS 2023 · 63 citations
