Predicting evolutionary rate as a pretraining task improves genome language model representations
Mica Consens, Kevin Yang, James Hall, Ashley Conard, BO WANG, Lorin Crawford, Alan Moses, Alex Lu
Abstract
Genome language models (gLM) have the potential to further understanding of regulatory genomics without requiring labeled data. Most gLMs are pretrained using sequence reconstruction tasks inspired by natural language processing, but recent studies have shown that these gLMs often fail to capture biological signal. To overcome this, we introduce pretraining tasks that predict the rate of evolution. These tasks are designed so that they can be composed with sequence reconstruction, enabling a controlled comparison of predicting sequence only, evolutionary rate only, or both. To address gaps in existing evaluations, we developed a suite of biologically grounded benchmarks. Across these tasks, and for established variant effect prediction benchmarks, models pretrained on both sequence and evolutionary rate outperform those trained on sequence alone, and training on evolutionary rate can make the even the relatively small models in our work competitive with much larger existing gLMs for some tasks on the human genome. These results establish evolution as a key training target for genome-scale models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08f2d257-a3d0-466b-b99b-2819fb53c433Builds on4
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao et al.ICML 2024 · 195 citations
- BEND: Benchmarking DNA Language Models on Biologically Meaningful TasksFrederikke Isa Marin, Felix Teufel, Marc Horlacher, Dennis Madsen et al.ICLR 2024 · 75 citations
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species GenomesZhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta et al.ICLR 2024 · 67 citations
- Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi et al.ICLR 2026 · 16 citations
Related papers
- GenomeQA: Benchmarking General Large Language Models for Genome Sequence UnderstandingWeicai Long, Yusen Hou, Junning Feng, Houcheng Su et al.ACL 2026
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- Interpreting Genomic Language Models using Sparse AutoencodersAkira Nair, Jaehyun Joo, Jonghyun Lee, Lina Takemaru et al.ICML 2026
- AntigenLM: Structure-Aware DNA Language Modeling for InfluenzaYue Pei, Xuebin Chi, Yu KangICLR 2026
- GENEB: Why Genomic Models Are Hard to CompareDaria Ledneva, Mikhail Nuridinov, Denis KuznetsovICML 2026
