From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language Models
Charles W. J. Pugh, Paulina G. Nuñez-Valencia, Mafalda Dias, Jonathan Frazer
摘要
Generative models trained on natural sequences are increasingly used to predict the effects of genetic variation, enabling progress in therapeutic design, disease risk prediction, and synthetic biology. In the zero-shot setting, variant impact is estimated by comparing the likelihoods of sequences, under the assumption that likelihood serves as a proxy for fitness. However, this assumption often breaks down in practice: sequence likelihood reflects not only evolutionary fitness constraints, but also phylogenetic structure and sampling biases, especially as model capacity increases. We introduce Likelihood-Fitness Bridging (LFB), a simple and general strategy that improves variant effect prediction by averaging model scores across sequences subject to similar selective pressures. Assuming an Ornstein-Uhlenbeck model of evolution, LFB can be viewed as a way to marginalize the effects of genetic drift, although its benefits appear to extend more broadly. LFB applies to existing protein and genomic language models without requiring retraining, and incurs only modest computational overhead. Evaluated on large-scale deep mutational scans and clinical benchmarks, LFB consistently improves predictive performance across model families and sizes. Notably, it reverses the performance plateau observed in larger protein language models, making the largest models the most accurate when combined with LFB. These results suggest that accounting for phylogenetic and sampling biases is essential to realizing the full potential of large sequence models in variant effect prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu 等NeurIPS 2021 · 被引用 969 次
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas 等NeurIPS 2023 · 被引用 574 次
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado 等ICML 2022 · 被引用 236 次
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 被引用 96 次
相关 Paper
- Protein Language Model Fitness is a Matter of PreferenceCade W. Gordon, Amy X. Lu, Pieter AbbeelICLR 2025
- Non-identifiability and the Blessings of Misspecification in Models of Molecular FitnessEli N. Weinstein, Alan Nawzad Amin, Jonathan Frazer, Debora S. MarksNeurIPS 2022 · 被引用 31 次
- Predicting evolutionary rate as a pretraining task improves genome language model representationsMica Consens, Kevin Yang, James Hall, Ashley Conard 等ICML 2026
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen 等ICML 2024 · 被引用 10 次
- A Variational Perspective on Generative Protein Fitness OptimizationLea Bogensperger, Dominik Narnhofer, Ahmed Allam, Konrad Schindler 等ICML 2025
