Scaling Laws and Architectural Frontiers in Metagenomic Foundation Models
Geraldene Munsamy, Gavin Ayres, Jérémie DONA, Carla Greco, Daniel P Anderson, Srijani Sridhar, William Chow, Aaron Kollasch, Robert Pecoraro, Tanggis Bohnuud, Keith Kam, Gus Minto-Cowcher
Abstract
Foundation models for genomics have the potential to revolutionize therapeutic design, yet the optimal architectural choices for modeling the vast and diverse distribution of metagenomic data remain under-explored. In this work, we present the machine learning methodology behind EDEN, a family of metagenomic foundation models scaled up to 28 billion parameters and trained on 9.7 trillion nucleotide tokens. We provide a systematic empirical study of architectural trade-offs between autoregressive Transformers (Llama-style), State-Space Models (Mamba), and Long-convolutional architectures (Hyena) for nucleotide-level modeling. Contrary to recent trends favoring linear-time sequence models for long-range biological data, we demonstrate that the Llama architecture exhibits superior scaling efficiency and semantic retrieval capabilities as the model capacity grows. We derive a set of quality-aware scaling laws for metagenomics, showing how model performance follows predictable power-law behavior across three orders of magnitude in parameters and data. Through extensive benchmarking, spanning unsupervised zeroshot fitness prediction, semantic completion, and gene recovery, we establish a blueprint for scaling biological foundation models and provide empirical evidence demonstrating why Transformer-based architectures define the current frontier.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7729d353-a0f2-4123-800c-c8b2c9c38cecBuilds on5
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao et al.ICML 2024 · 195 citations
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species GenomesZhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta et al.ICLR 2024 · 67 citations
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran et al.NeurIPS 2025 · 63 citations
Related papers
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong et al.NeurIPS 2024 · 44 citations
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationZehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li et al.NeurIPS 2025 · 6 citations
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann et al.NeurIPS 2025 · 8 citations
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthXuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen et al.NeurIPS 2024 · 63 citations
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour et al.ICML 2026
