Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA
Lifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai, Chaoqi Liang, Xinzhu Ma, Nanqing Dong, Wanli Ouyang
摘要
Foundation models have made significant strides in understanding the genomic language of DNA sequences. However, previous models typically adopt the tokenization methods designed for natural language, which are unsuitable for DNA sequences due to their unique characteristics. In addition, the optimal approach to tokenize DNA remains largely under-explored, and may not be intuitively understood by humans even if discovered. To address these challenges, we introduce MxDNA, a novel framework where the model autonomously learns an effective DNA tokenization strategy through gradient decent. MxDNA employs a sparse Mixture of Convolution Experts coupled with a deformable convolution to model the tokenization process, with the discontinuous, overlapping, and ambiguous nature of meaningful genomic segments explicitly considered. On Nucleotide Transformer Benchmarks and Genomic Benchmarks, MxDNA demonstrates superior performance to existing methods with less pretraining data and time, highlighting its effectiveness. Finally, we show that MxDNA learns unique tokenization strategy distinct to those of previous methods and captures genomic functionalities at a token level during self-supervised pretraining. Our MxDNA aims to provide a new perspective on DNA tokenization, potentially offering broad applications in various domains and yielding profound insights.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi 等ICLR 2026 · 被引用 16 次
- Transducing Language ModelsVésteinn Snæbjarnarson, Samuel Kiegeland, Tianyu Liu, Reda Boumasmoud 等ICLR 2026 · 被引用 3 次
- MergeDNA: Context-Aware Genome Modeling with Dynamic Tokenization Through Token MergingSiyuan Li, Kai Yu, Anna Wang, Zicheng Liu 等AAAI 2026 · 被引用 2 次
- PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNAAlice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M Athar, Agnieszka Dobrowolska 等ICLR 2026 · 被引用 2 次
- Multimodal Medical Code TokenizerXiaorui Su, Shvat Messica, Yepeng Huang, Ruth Johnson 等ICML 2025 · 被引用 2 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
相关 Paper
- dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence LearningArnav Shah, Junzhe Li, Parsa Idehpour, Adibvafa Fallahpour 等ICML 2026
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable RepresentationsKe Ding, Brian J. Parker, Jiayu WenAAAI 2026 · 被引用 1 次
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsTaewon Kim, Jihwan Shin, Hyomin Kim, Youngmok Jung 等ICML 2026 · 被引用 1 次
- LDARNet: DNA Adaptive Representation Network with Learnable Tokenization for Genomic ModelingDaria Ledneva, Denis KuznetsovICML 2026 · 被引用 1 次
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann 等NeurIPS 2025 · 被引用 8 次
