Scaling Laws and Architectural Frontiers in Metagenomic Foundation Models
Geraldene Munsamy, Gavin Ayres, Jérémie DONA, Carla Greco, Daniel P Anderson, Srijani Sridhar, William Chow, Aaron Kollasch, Robert Pecoraro, Tanggis Bohnuud, Keith Kam, Gus Minto-Cowcher
摘要
Foundation models for genomics have the potential to revolutionize therapeutic design, yet the optimal architectural choices for modeling the vast and diverse distribution of metagenomic data remain under-explored. In this work, we present the machine learning methodology behind EDEN, a family of metagenomic foundation models scaled up to 28 billion parameters and trained on 9.7 trillion nucleotide tokens. We provide a systematic empirical study of architectural trade-offs between autoregressive Transformers (Llama-style), State-Space Models (Mamba), and Long-convolutional architectures (Hyena) for nucleotide-level modeling. Contrary to recent trends favoring linear-time sequence models for long-range biological data, we demonstrate that the Llama architecture exhibits superior scaling efficiency and semantic retrieval capabilities as the model capacity grows. We derive a set of quality-aware scaling laws for metagenomics, showing how model performance follows predictable power-law behavior across three orders of magnitude in parameters and data. Through extensive benchmarking, spanning unsupervised zeroshot fitness prediction, semantic completion, and gene recovery, we establish a blueprint for scaling biological foundation models and provide empirical evidence demonstrating why Transformer-based architectures define the current frontier.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas 等NeurIPS 2023 · 被引用 574 次
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species GenomesZhihan Zhou, Yanrong Ji, Weijian Li, Pratik Dutta 等ICLR 2024 · 被引用 67 次
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran 等NeurIPS 2025 · 被引用 63 次
相关 Paper
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong 等NeurIPS 2024 · 被引用 44 次
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationZehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li 等NeurIPS 2025 · 被引用 6 次
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann 等NeurIPS 2025 · 被引用 8 次
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context LengthXuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen 等NeurIPS 2024 · 被引用 63 次
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour 等ICML 2026
