dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning
Arnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour, Brandon Wang, Sukjun Hwang, BO WANG, Patrick Hsu, Hani Goodarzi, Albert Gu
Abstract
Genomic foundation models have the potential to decode DNA syntax, yet face a fundamental tradeoff in their input representation. Standard fixed-vocabulary tokenizers fragment biologically meaningful motifs such as codons and regulatory elements, while nucleotide-level models preserve biological coherence but incur prohibitive computational costs for long contexts. We introduce dnaHNet, a state-of-the-art tokenizer-free autoregressive model that segments and models genomic sequences end-to-end. Using a differentiable dynamic chunking mechanism, dnaHNet compresses raw nucleotides into latent tokens adaptively, balancing compression with predictive accuracy. Pretrained on prokaryotic genomes, dnaHNet outperforms leading architectures including StripedHyena2 in scaling and efficiency. This recursive chunking yields quadratic FLOP reductions, enabling > 3× inference speedup over Transformers. On zero-shot tasks, dnaH-Net achieves superior performance in predicting protein variant fitness and gene essentiality, while automatically discovering hierarchical biological structures without supervision. These results establish dnaHNet as a scalable, interpretable framework for next-generation genomic modeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0445ba4-d1b1-4e4b-b293-66f92695e509Builds on8
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- Byte Latent Transformer: Patches Scale Better Than TokensArtidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez, John Nguyen et al.ACL 2025 · 116 citations
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 76 citations
Related papers
- LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic ModelingDaria Ledneva, Denis KuznetsovICML 2026 · 1 citation
- BioToken and BioFM – Biologically-Informed Tokenization Enables Accurate and Efficient Genomic Foundation ModelsAleksandr Medvedev, Karthik Viswanathan, Praveenkumar Kanithi, Kirill Vishniakov et al.ICML 2026
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai et al.NeurIPS 2024 · 23 citations
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu et al.AAAI 2026 · 2 citations
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsTaewon Kim, Jihwan Shin, Hyomin Kim, Youngmok Jung et al.ICML 2026 · 1 citation
