Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual Annotation
Zehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li, Wentao Gu, Guy-Bart Stan, Yiren Zhao, Caihua Shan
Abstract
The interpretation of genomic sequences is crucial for understanding biological processes. To handle the growing volume of DNA sequence data, Genomic Foundation Models (GFMs) have been developed by adapting architectures and training paradigms from Large Language Models (LLMs). Despite their remarkable performance in DNA sequence classification tasks, there remains a lack of systematic understanding regarding the pre-training and task-adaptation processes of GFMs. Moreover, existing GFMs cannot achieve state-of-the-art performance on both short and long-context tasks and lack multimodal abilities. By revisiting pre-training architectures and post-training techniques, we propose OMNI-DNA, a family of models spanning 20M to 1.1B parameters that supports sequence understanding, long-context genomic reasoning, and natural-language annotation. Omni-DNA establishes new state-of-the-art results on 18 of 26 evaluations drawn from Nucleotide Transformer and Genomic Benchmarks. When jointly finetuning on biologically related tasks, Omni-DNA consistently outperforms existing models and demonstrates multi-tasking abilities. Furthermore, we introduce SEQ-PACK, an adaptive compression mechanism that enables efficient long-context modeling by summarizing historical tokens through position-aware learnable sampling. This allows transformer-based models to process ultra-long genomic sequences with minimal memory and computational overhead. Leveraging SEQPACK, Omni-DNA excels at enhancer-target interaction prediction, capturing distal regulatory effects over 450kbp. Finally, we present SEQ2FUNC, a newly constructed dataset that empowers Omni-DNA to generate accurate and functionally meaningful interpretations of DNA sequences, opening new avenues for genomic analysis and discovery.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas et al.NeurIPS 2023 · 574 citations
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao et al.ICML 2024 · 195 citations
- NEFTune: Noisy Embeddings Improve Instruction FinetuningNeel Jain, Ping-yeh Chiang, Yuxin Wen, John Kirchenbauer et al.ICLR 2024 · 120 citations
Related papers
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann et al.NeurIPS 2025 · 8 citations
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsWeimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin et al.ICML 2026
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable RepresentationsKe Ding, Brian J. Parker, Jiayu WenAAAI 2026 · 1 citation
- Rethinking Genomic Modeling Through Optical Character RecognitionHongxin Xiang, Pengsen Ma, Yunkang Cao, Di Yu et al.ICML 2026
- Scaling Laws and Architectural Frontiers in Metagenomic Foundation ModelsGeraldene Munsamy, Gavin Ayres, Jérémie DONA, Carla Greco et al.ICML 2026
