Frozen Pretrained Transformers as Universal Computation Engines
Kevin Lu, Aditya Grover, Pieter Abbeel, Igor Mordatch
Abstract
We investigate the capability of a transformer pretrained on natural language to generalize to other modalities with minimal finetuning -in particular, without finetuning of the self-attention and feedforward layers of the residual blocks. We consider such a model, which we call a Frozen Pretrained Transformer (FPT), and study finetuning it on a variety of sequence classification tasks spanning numerical computation, vision, and protein fold prediction. In contrast to prior works which investigate finetuning on the same modality as the pretraining dataset, we show that pretraining on natural language can improve performance and compute efficiency on non-language downstream tasks. Additionally, we perform an analysis of the architecture, comparing the performance of a random initialized transformer to a random LSTM. Combining the two insights, we find language-pretrained transformers can obtain strong performance on a variety of non-language tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e106617d-40fe-416d-aaad-e9489f018627Cited by top-tier papers25
- One Fits All: Power General Time Series Analysis by Pretrained LMTian Zhou, Peisong Niu, Xue Wang, Liang Sun et al.NeurIPS 2023 · 1,178 citations
- MOMENT: A Family of Open Time-series Foundation ModelsMononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai et al.ICML 2024 · 442 citations
- VeRA: Vector-based Random Matrix AdaptationDawid Jan Kopiczko, Tijmen Blankevoort, Yuki M. AsanoICLR 2024 · 308 citations
- Generating Automatic Feedback on UI Mockups with Large Language ModelsPeitong Duan, Jeremy Warner, Yang Li, Bjoern HartmannCHI 2024 · 81 citations
- Cross-Modal Fine-Tuning: Align then RefineJunhong Shen, Liam Li, Lucio M. Dery, Corey Staten et al.ICML 2023 · 64 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Understanding Transfer Learning of RNA Foundation Models on Downstream TasksYuan Li, Heng Yang, Renzhi Chen, Ke LiICML 2026
- Transformer Layers as PaintersQi Sun, Marc Pickett, Aakash Kumar Nain, Llion JonesAAAI 2025 · 49 citations
- Vision Transformers are Parameter-Efficient Audio-Visual LearnersYan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal et al.CVPR 2023
- Feature Reuse and Scaling: Understanding Transfer Learning with Protein Language ModelsFrancesca-Zhoufan Li, Ava P. Amini, Yisong Yue, Kevin K. Yang et al.ICML 2024 · 61 citations
- Frozen Transformers in Language Models Are Effective Visual Encoder LayersZiqi Pang, Ziyang Xie, Yunze Man, Yu-Xiong WangICLR 2024 · 54 citations
