Evolutionary Reasoning Does Not Arise in Standard Usage of Protein Language Models
Yasha Ektefaie, Andrew Shen, Lavik Jain, Maha R. Farhat, Marinka Zitnik
Abstract
Protein language models (PLMs) are often assumed to capture evolutionary information by training on large protein sequence datasets. Yet it remains unclear whether PLMs can reason about evolution-that is, infer evolutionary relationships between sequences. We test this capability by evaluating whether standard PLM usage, frozen or fine-tuned embeddings with distance-based comparison, supports evolutionary reasoning. Existing PLMs consistently fail to recover phylogenetic structure, despite strong performance on sequence-level tasks such as masked-token and contact prediction. We present PHYLA, a hybrid state-space and transformer model that jointly processes multiple sequences and is trained using a tree-based objective across 3,000 phylogenies spanning diverse protein families. PHYLA outperforms the next-best PLM by 9% on tree reconstruction and 23% on taxonomic clustering while remaining alignment-and guide-tree-free. Although classical alignment pipelines achieve higher absolute accuracy, PHYLA narrows the gap and achieves markedly lower end-to-end runtime. Applied to real data, PHYLA reconstructs biologically accurate clades in the tree of life and resolves genome-scale relationships among Mycobacterium tuberculosis isolates. These findings suggest that, under standard usage, evolutionary reasoning does not reliably emerge from large-scale sequence modeling. Instead, PHYLA shows that models trained with phylogenetic supervision can reason about evolution more effectively, offering a biologically grounded path toward evolutionary foundation models.
39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11381735-b380-4f62-997c-eae63857b566Builds on11
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier et al.ICML 2021 · 686 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
Related papers
- Nested birth-death processes are competitive with neural networks as time-dependent models of protein evolutionAnnabel Large, Ian HolmesICML 2026
- AntigenLM: Structure-Aware DNA Language Modeling for InfluenzaYue Pei, Xuebin Chi, Yu KangICLR 2026
- FlexRibbon: Joint Sequence and Structure Pretraining for Protein ModelingJianwei Zhu, Yu Shi, Ran Bi, Peiran Jin et al.ICLR 2026 · 2 citations
- Distilling Structural Representations into Protein Sequence ModelsJeffrey Ouyang-Zhang, Chengyue Gong, Yue Zhao, Philipp Krähenbühl et al.ICLR 2025
- DPLM-2: A Multimodal Diffusion Protein Language ModelXinyou Wang, Zaixiang Zheng, Fei Ye, Dongyu Xue et al.ICLR 2025
