Scaling Unlocks Broader Generation and Deeper Functional Understanding of Proteins
Aadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran, Alexander Hoffnagle, Kyle Ching, Michael Martyn, Stephen Nayfach, Jeffrey A. Ruffolo, Ali Madani
Abstract
Generative protein language models (PLMs) are powerful tools for designing proteins purpose-built to solve problems in medicine, agriculture, and industrial processes. Recent work has trained ever larger language models, but there has been little systematic study of the optimal training distributions and the influence of model scale on the sequences generated by PLMs. We introduce the ProGen3 family of sparse generative PLMs, and we develop compute-optimal scaling laws to scale up to a 46B-parameter model pre-trained on 1.5T amino acid tokens. Pro-Gen3's pre-training data is sampled from an optimized data distribution over the Profluent Protein Atlas v1, a carefully curated dataset of 3.4B full-length proteins. We evaluate for the first time in the wet lab the influence of model scale on the sequences generated by PLMs, and we find that larger models generate viable proteins for a much wider diversity of protein families. Finally, we find both computationally and experimentally that larger models are more responsive to alignment with laboratory data, resulting in improved protein fitness prediction and sequence generation capabilities. These results indicate that larger PLMs like ProGen3-46B trained on larger, well-curated datasets are powerful foundation models that push the frontier of protein design. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15b197cb-7518-45fa-90da-4abf979f0299Cited by top-tier papers7
- Understanding protein function with a multimodal retrieval-augmented foundation modelTimothy F. Truong Jr., Tristan BeplerNeurIPS 2025 · 15 citations
- From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language ModelsCharles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan FrazerNeurIPS 2025 · 12 citations
- Protein Inverse Folding From Structure FeedbackJunde Xu, Zijun Gao, Xinyi Zhou, Jie Hu et al.NeurIPS 2025 · 9 citations
- Protein Circuit Tracing via Cross-layer TranscodersDarin Tsui, Kunal Talreja, Daniel Saeedi, Amirali AghazadehICML 2026 · 4 citations
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone GenerationChao Song, Zhiyuan Liu, Han Huang, Liang Wang et al.NeurIPS 2025 · 3 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong et al.NeurIPS 2024 · 44 citations
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 96 citations
- Concept Bottleneck Language Models For Protein DesignAya Abdelsalam Ismail, Tuomas P. Oikarinen, Amy Wang, Julius Adebayo et al.ICLR 2025
- Diffusion Language Models Are Versatile Protein LearnersXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICML 2024 · 113 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
