PoET: A generative model of protein families as sequences-of-sequences
Timothy F. Truong Jr., Tristan Bepler
Abstract
Generative protein language models are a natural way to design new proteins with desired functions. However, current models are either difficult to direct to produce a protein from a specific family of interest, or must be trained on a large multiple sequence alignment (MSA) from the specific family of interest, making them unable to benefit from transfer learning across families. To address this, we propose rtein volutionary ransformer (PoET), an autoregressive generative model of whole protein families that learns to generate sets of related proteins as sequences-of-sequences across tens of millions of natural protein sequence clusters. PoET can be used as a retrieval-augmented language model to generate and score arbitrary modifications conditioned on any protein family of interest, and can extrapolate from short context lengths to generalize well even for small families. This is enabled by a unique Transformer layer; we model tokens sequentially within sequences while attending between sequences order invariantly, allowing PoET to scale to context lengths beyond those used during training. In extensive experiments on deep mutational scanning datasets, we show that PoET outperforms existing protein language models and evolutionary sequence models for variant function prediction across proteins of all MSA depths. We also demonstrate PoET's ability to controllably generate new protein sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6de98f7e-b0a9-4fb5-890c-4924f187ca1eCited by top-tier papers20
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran et al.NeurIPS 2025 · 63 citations
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li et al.NeurIPS 2024 · 20 citations
- Multi-Scale Representation Learning for Protein Fitness PredictionZuobai Zhang, Pascal Notin, Yining Huang, Aurélie C. Lozano et al.NeurIPS 2024 · 17 citations
- Understanding protein function with a multimodal retrieval-augmented foundation modelTimothy F. Truong Jr., Tristan BeplerNeurIPS 2025 · 15 citations
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 12 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin et al.EMNLP 2024 · 4 citations
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong et al.NeurIPS 2024 · 44 citations
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- ProSE: the architecture and design of a protein discovery engineEyes Robson, Ceyu Xu, Lisa Wu WillsASPLOS 2022 · 9 citations
