Magneto: A Foundation Transformer
Hongyu Wang, Shuming Ma, Shaohan Huang, Li Dong, Wenhui Wang, Zhiliang Peng, Yu Wu, Payal Bajaj, Saksham Singhal, Alon Benhaim, Barun Patra, Zhun Liu
Abstract
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different implementations for better performance, e.g., Post-LayerNorm for BERT, and Pre-LayerNorm for GPT and vision Transformers. We call for the development of Foundation Transformer for true general-purpose modeling, which serves as a goto architecture for various tasks and modalities with guaranteed training stability. In this work, we introduce a Transformer variant, named MAG-NETO, to fulfill the goal. Specifically, we propose Sub-LayerNorm for good expressivity, and the initialization strategy theoretically derived from DeepNet (Wang et al., 2022a) for stable scaling up. Extensive experiments demonstrate its superior performance and better stability than the de facto Transformer variants designed for various applications, including language modeling (i.e., BERT, and GPT), machine translation, vision pretraining (i.e., BEiT), speech recognition, and multimodal pretraining (i.e., BEiT-3).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75920a25-5df5-4f09-a9da-6987206b2ed7Cited by top-tier papers4
- You Only Cache Once: Decoder-Decoder Architectures for Language ModelsYutao Sun, Li Dong, Yi Zhu, Shaohan Huang et al.NeurIPS 2024 · 162 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
- Lr0.Fm: low-Resolution Zero-Shot Classification Benchmark for Foundation ModelsPriyank Pathak, Shyam Marjit, Shruti Vyas, Yogesh S. RawatICLR 2025
- Differential TransformerTianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun et al.ICLR 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language TasksWenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck et al.CVPR 2023
- Pay Attention to MLPsHanxiao Liu, Zihang Dai, David R. So, Quoc V. LeNeurIPS 2021 · 912 citations
- Frozen Pretrained Transformers as Universal Computation EnginesKevin Lu, Aditya Grover, Pieter Abbeel, Igor MordatchAAAI 2022 · 133 citations
- MatFormer: Nested Transformer for Elastic InferenceDevvrit, Sneha Kudugunta, Aditya Kusupati, Tim Dettmers et al.NeurIPS 2024 · 97 citations
- Deep Compression of Pre-trained Transformer ModelsNaigang Wang, Chi-Chun (Charlie) Liu, Swagath Venkataramani, Sanchari Sen et al.NeurIPS 2022 · 38 citations
