Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?
Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi, Marco A. F. Pimentel, Ronnie Rajan, Shadab Khan
Abstract
The success of Large Language Models has inspired the development of Genomic Foundation Models (GFMs) through similar pretraining techniques. However, the relationship between pretraining performance and effectiveness in downstream genomic tasks remains unclear. Additionally, the high computational cost of pretraining raises questions about its cost-efficiency. To assess the usefulness of pretraining in genomics, we evaluated seven different GFMs across 52 diverse genomic tasks, comparing them to their counterparts with randomly initialized weights. Across benchmarks, we find that randomly initialized models provide surprisingly strong baselines and tokenizer and architecture choices strongly shape both these baselines and the gains from pretraining. Specifically, character-token models often match or exceed the performance of larger pretrained k-mer or BPE models, whereas subword models appear to benefit from pretraining. We also find that the evaluated GFMs fail to capture clinically relevant genetic mutations, with embeddings and log-likelihood ratios showing limited sensitivity to annotated variants. For the tasks we study, these results suggest that current NLP-style pretraining strategies provide modest, tokenizer-gated improvements over strong random baselines and motivate more biologically informed tokenization and variant-aware objectives. Our code is available at github.com/m42- health/gfm-random-eval.
- S.K. was with M42 (initially) and ADIA Lab (subsequently) during this research work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f0b7da1-1e33-431c-9da9-97c6db23d575Cited by top-tier papers3
- Predicting evolutionary rate as a pretraining task improves genome language model representationsMica Consens, Kevin Yang, James Hall, Ashley Conard et al.ICML 2026
- BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation ModelsAleksandr Medvedev, Karthik Viswanathan, Praveenkumar Kanithi, Kirill Vishniakov et al.ICML 2026
- BIOARC: Discovering Optimal Neural Architectures for Biological Foundation ModelsYi Fang, Haoran Xu, Jiaxin Han, Sirui Ding et al.ICML 2026
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao et al.ICML 2024 · 195 citations
Related papers
- GENEB: Why Genomic Models Are Hard to CompareDaria Ledneva, Mikhail Nuridinov, Denis KuznetsovICML 2026
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species GenomesZhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta et al.ICLR 2024 · 67 citations
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsWeimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin et al.ICML 2026
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu et al.AAAI 2026 · 2 citations
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable RepresentationsKe Ding, Brian J. Parker, Jiayu WenAAAI 2026 · 1 citation
