BEND: Benchmarking DNA Language Models on Biologically Meaningful Tasks
Frederikke Isa Marin, Felix Teufel, Marc Horlacher, Dennis Madsen, Dennis Pultz, Ole Winther, Wouter Boomsma
摘要
The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence remains both expensive and challenging. This has sparked interest in unsupervised language modeling of genomic DNA, a paradigm that has seen great success for protein sequence data. Although various DNA language models have been proposed, evaluation tasks often differ between individual works, and might not fully recapitulate the fundamental challenges of genome annotation, including the length, scale and sparsity of the data. In this study, we introduce BEND, a Benchmark for DNA language models, featuring a collection of realistic and biologically meaningful downstream tasks defined on the human genome. We find that embeddings from current DNA LMs can approach performance of expert methods on some tasks, but only capture limited information about long-range features. BEND is available at https://github.com/frederikkemarin/BEND .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- BioReason: Incentivizing Multimodal Biological Reasoning within a DNA-LLM ModelAdibvafa Fallahpour, Andrew Magnuson, Purav Gupta, Shihao Ma 等NeurIPS 2025 · 被引用 53 次
- Tokenization to Transfer: Do Genomic Foundation Models Learn Good Representations?Kirill Vishniakov, Karthik Viswanathan, Aleksandr Medvedev, Praveenkumar Kanithi 等ICLR 2026 · 被引用 16 次
- PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNAAlice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M Athar, Agnieszka Dobrowolska 等ICLR 2026 · 被引用 2 次
- DNACHUNKER: Learnable Tokenization for DNA Language ModelsTaewon Kim, Jihwan Shin, Hyomin Kim, Youngmok Jung 等ICML 2026 · 被引用 1 次
- Predicting evolutionary rate as a pretraining task improves genome language model representationsMica Consens, Kevin Yang, James Hall, Ashley Conard 等ICML 2026
它引用的顶会 Paper5
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas 等NeurIPS 2023 · 被引用 574 次
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu 等ICML 2023 · 被引用 481 次
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov 等ICLR 2021 · 被引用 366 次
- BERTology Meets Biology: Interpreting Attention in Protein Language ModelsJesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong 等ICLR 2021 · 被引用 357 次
相关 Paper
- GenomeQA: Benchmarking General Large Language Models for Genome Sequence UnderstandingWeicai Long, Yusen Hou, Junning Feng, Houcheng Su 等ACL 2026
- Genome-Factory: A Library for Tuning, Deploying, and Interpreting Genomic Foundation ModelsWeimin Wu, Xuefeng Song, Yibo Wen, Qinjie Lin 等ICML 2026
- JanusDNA: A Powerful Bi-directional Hybrid DNA Foundation ModelQihao Duan, Bingding Huang, Zhenqiao Song, Irina Lehmann 等NeurIPS 2025 · 被引用 8 次
- Hyperbolic Genome EmbeddingsRaiyan R. Khan, Philippe Chlenski, Itsik Pe'erICLR 2025
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationZehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li 等NeurIPS 2025 · 被引用 6 次
