PoET: A generative model of protein families as sequences-of-sequences
Timothy F. Truong Jr., Tristan Bepler
摘要
Generative protein language models are a natural way to design new proteins with desired functions. However, current models are either difficult to direct to produce a protein from a specific family of interest, or must be trained on a large multiple sequence alignment (MSA) from the specific family of interest, making them unable to benefit from transfer learning across families. To address this, we propose rtein volutionary ransformer (PoET), an autoregressive generative model of whole protein families that learns to generate sets of related proteins as sequences-of-sequences across tens of millions of natural protein sequence clusters. PoET can be used as a retrieval-augmented language model to generate and score arbitrary modifications conditioned on any protein family of interest, and can extrapolate from short context lengths to generalize well even for small families. This is enabled by a unique Transformer layer; we model tokens sequentially within sequences while attending between sequences order invariantly, allowing PoET to scale to context lengths beyond those used during training. In extensive experiments on deep mutational scanning datasets, we show that PoET outperforms existing protein language models and evolutionary sequence models for variant function prediction across proteins of all MSA depths. We also demonstrate PoET's ability to controllably generate new protein sequences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran 等NeurIPS 2025 · 被引用 63 次
- MSAGPT: Neural Prompting Protein Structure Prediction via MSA Generative Pre-TrainingBo Chen, Zhilei Bei, Xingyi Cheng, Pan Li 等NeurIPS 2024 · 被引用 20 次
- Multi-Scale Representation Learning for Protein Fitness PredictionZuobai Zhang, Pascal Notin, Yining Huang, Aurélie C. Lozano 等NeurIPS 2024 · 被引用 17 次
- Understanding protein function with a multimodal retrieval-augmented foundation modelTimothy F. Truong Jr., Tristan BeplerNeurIPS 2025 · 被引用 15 次
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 被引用 1,246 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
相关 Paper
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
- Retrieved Sequence Augmentation for Protein Representation LearningChang Ma, Haiteng Zhao, Lin Zheng, Jiayi Xin 等EMNLP 2024 · 被引用 4 次
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong 等NeurIPS 2024 · 被引用 44 次
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue 等ICML 2024 · 被引用 113 次
- ProSE: the architecture and design of a protein discovery engineEyes Robson, Ceyu Xu, Lisa Wu WillsASPLOS 2022 · 被引用 9 次
