Ultrafast classical phylogenetic method beats large protein language models on variant effect prediction
Sebastian Prillo, Wilson Wu, Yun Song
Abstract
Amino acid substitution rate matrices are fundamental to statistical phylogenetics and evolutionary biology. Estimating them typically requires reconstructed trees for massive amounts of aligned proteins, which poses a major computational bottleneck. In this paper, we develop a near-linear time method to estimate these rate matrices from multiple sequence alignments (MSAs) alone, thereby speeding up computation by orders of magnitude. Our method relies on a near-linear time cherry reconstruction algorithm which we call FastCherries and it can be easily applied to MSAs with millions of sequences. On both simulated and real data, we demonstrate the speed and accuracy of our method as applied to the classical model of protein evolution. By leveraging the unprecedented scalability of our method, we develop a new, rich phylogenetic model called SiteRM, which can estimate a general site-specific rate matrix for each column of an MSA. Remarkably, in variant effect prediction for both clinical and deep mutational scanning data in ProteinGym, we show that despite being an independent-sites model, our SiteRM model outperforms large protein language models that learn complex residue-residue interactions between different sites. We attribute our increased performance to conceptual advances in our probabilistic treatment of evolutionary data and our ability to handle extremely large MSAs. We anticipate that our work will have a lasting impact across both statistical phylogenetics and computational variant effect prediction. FastCherries and SiteRM are implemented in the CherryML package https://github.com/songlab-cal/CherryML.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Conditionally Site-Independent Neural Evolution of Antibody SequencesStephen Lu, Aakarsh Vermani, Kohei Sanno, Jiarui Lu et al.ICML 2026 · 2 citations
- Nested birth-death processes are competitive with neural networks as time-dependent models of protein evolutionAnnabel Large, Ian HolmesICML 2026
Builds on4
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier et al.ICML 2021 · 686 citations
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin et al.ICML 2022 · 560 citations
- Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference-time RetrievalPascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena-Hurtado et al.ICML 2022 · 236 citations
Related papers
- Predicting evolutionary rate as a pretraining task improves genome language model representationsMica Consens, Kevin Yang, James Hall, Ashley Conard et al.ICML 2026
- Fast and More Powerful Selective Inference for Sparse High-Order Interaction ModelDiptesh Das, Vo Nguyen Le Duy, Hiroyuki Hanada, Koji Tsuda et al.AAAI 2022 · 23 citations
- PoET: A generative model of protein families as sequences-of-sequencesTimothy F. Truong Jr., Tristan BeplerNeurIPS 2023 · 96 citations
- Evolution-Inspired Loss Functions for Protein Representation LearningChengyue Gong, Adam R. Klivans, James Loy, Tianlong Chen et al.ICML 2024 · 10 citations
- DePLM: Denoising Protein Language Models for Property OptimizationZeyuan Wang, Keyan Ding, Ming Qin, Xiaotong Li et al.NeurIPS 2024 · 2 citations
