The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling
Andre Cornman, Jacob West-Roberts, Antonio Pedro Camargo, Simon Roux, Martin Beracochea, Milot Mirdita, Sergey Ovchinnikov, Yunha Hwang
Abstract
Abstract Biological language model performance depends heavily on pretraining data quality, diversity, and size. While metagenomic datasets feature enormous biological diversity, their utilization as pretraining data has been limited due to challenges in data accessibility, quality filtering and deduplication. Here, we present the Open MetaGenomic (OMG) corpus, a genomic pretraining dataset totalling 3.1T base pairs and 3.3B protein coding sequences, obtained by combining two largest metagenomic dataset repositories (JGI’s IMG and EMBL’s MGnify). We first document the composition of the dataset and describe the quality filtering steps taken to remove poor quality data. We make the OMG corpus available as a mixed-modality genomic sequence dataset that represents multi-gene encoding genomic sequences with translated amino acids for protein coding sequences, and nucleic acids for intergenic sequences. We train the first mixed-modality genomic language model (gLM2) that leverages genomic context information to learn robust functional representations, as well as coevolutionary signals in protein-protein interfaces and genomic regulatory syntax. Furthermore, we show that deduplication in embedding space can be used to balance the corpus, demonstrating improved performance on downstream tasks. The OMG dataset is publicly hosted on the Hugging Face Hub at https://huggingface.co/datasets/tattabio/OMG and gLM2 is available at https://huggingface.co/tattabio/gLM2_650M .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3982074e-8cc0-414c-857a-96fed4b45cc8Cited by top-tier papers2
- Scaling Unlocks Broader Generation and Deeper Functional Understanding of ProteinsAadyot Bhatnagar, Sarthak Jain, Joel Beazer, Samuel Curran et al.NeurIPS 2025 · 63 citations
- On the Relationship Between Activation Outliers and Feature Death in Sparse AutoencodersElana Simon, Etowah Adams, James ZouICML 2026
Builds on3
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled CorpusJesse Dodge, Maarten Sap, Ana Marasovic, William Agnew et al.EMNLP 2021 · 18 citations
Related papers
- Training Compute-Optimal Protein Language ModelsXingyi Cheng, Bo Chen, Pan Li, Jing Gong et al.NeurIPS 2024 · 44 citations
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-TrainingPierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza et al.ICLR 2026 · 22 citations
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationZehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li et al.NeurIPS 2025 · 6 citations
- OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with TextQingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang et al.ICLR 2025
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann et al.NeurIPS 2025 · 8 citations
