LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic Modeling
Daria Ledneva, Denis Kuznetsov
摘要
Genomic foundation models increasingly adopt large language model architectures, yet almost universally rely on fixed tokenization schemes such as -mers, BPE, or single nucleotides, which impose arbitrary sequence boundaries that may obscure biologically relevant structure. We present LDARNet, a 120M-parameter hierarchical genomic foundation model that adapts H-Net-style dynamic chunking from autoregressive generation to masked language modeling, combining BiMamba-2 state-space layers with local attention, bidirectional routing, and a ratio-based regularizer to induce adaptive token boundaries without supervision. Fine-tuned on 27 tasks from the Nucleotide Transformer and Genomic Benchmarks suites, LDARNet achieves 11/18 wins among compact models (300M parameters) and state-of-the-art results on 5 histone modification tasks, outperforming models up to 20 larger. A FLOPs-matched controlled experiment isolates learned routing as the source of these gains: learned boundaries beat fixed-grid boundaries by up to 14 percentage points on histone tasks at identical compute. Nucleotide-resolution analysis further shows that the learned boundaries align with canonical promoter motifs and splice junctions without supervision, providing a biological interpretation for adaptive tokenization in genomic foundation models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- Dynamic Chunking for End-to-End Hierarchical Sequence ModelingSukjun Hwang, Brandon Wang, Albert GuICLR 2026 · 被引用 76 次
相关 Paper
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour 等ICML 2026
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable RepresentationsKe Ding, Brian J. Parker, Jiayu WenAAAI 2026 · 被引用 1 次
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsTaewon Kim, Jihwan Shin, Hyomin Kim, Youngmok Jung 等ICML 2026 · 被引用 1 次
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu 等AAAI 2026 · 被引用 2 次
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai 等NeurIPS 2024 · 被引用 23 次
