Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis
Mohammadsaleh Refahi, Mahdi Abavisani, Bahrad A. Sokhansanj, James R. Brown, Gail Rosen
摘要
Transformers have revolutionized nucleotide sequence analysis, yet capturing long-range dependencies remains challenging. Recent studies show that autoregressive transformers often exhibit Markovian behavior by relying on fixed-length context windows for next-token prediction. However, standard self-attention mechanisms are computationally inefficient for long sequences due to their quadratic complexity and do not explicitly enforce global transition consistency. We introduce CARMANIA (Context-Aware Regularization with Markovian Integration for Attention-Based Nucleotide Analysis), a self-supervised pretraining framework that augments next-token (NT) prediction with a transition-matrix (TM) loss. The TM loss aligns predicted token transitions with empirically derived ngram statistics from each input sequence, encouraging the model to capture higherorder dependencies beyond local context. This integration enables CARMANIA to learn organism-specific sequence structures that reflect both evolutionary constraints and functional organization. We evaluate CARMANIA across diverse genomic tasks, including regulatory element prediction, functional gene classification, taxonomic inference, antimicrobial resistance detection, and biosynthetic gene cluster classification. CARMANIA outperforms the previous best long-context model by at least 7%, matches state-ofthe-art on shorter sequences (exceeding prior results on 20/40 tasks while running ∼2.5× faster), and shows particularly strong improvements on enhancer and housekeeping gene classification tasks-including up to a 34% absolute gain in Matthews correlation coefficient (MCC) for enhancer prediction. The TM loss boosts accuracy in 33 of 40 tasks, especially where local motifs or regulatory patterns drive prediction. This enables more effective modeling of sequence-dependent biological features while maintaining robustness across non-coding and low-signal regions. Code available at https://github.com/EESI/carmania.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide ResolutionEric Nguyen, Michael Poli, Marjan Faizi, Armin W. Thomas 等NeurIPS 2023 · 被引用 574 次
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach 等NeurIPS 2024 · 被引用 140 次
相关 Paper
- NucEL: Single-Nucleotide ELECTRA-Style Genomic Pre-training for Efficient and Interpretable RepresentationsKe Ding, Brian J. Parker, Jiayu WenAAAI 2026 · 被引用 1 次
- Curriculum Learning for Biological Sequence Prediction: The Case of De Novo Peptide SequencingXiang Zhang, Jiaqi Wei, Zijie Qiu, Sheng Xu 等ICML 2025
- Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNALifeng Qiao, Peng Ye, Yuchen Ren, Weiqiang Bai 等NeurIPS 2024 · 被引用 23 次
- Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven PriorsIdo Amos, Jonathan Berant, Ankit GuptaICLR 2024 · 被引用 39 次
- Omni-DNA: A Genomic Model Supporting Sequence Understanding, Long-context, and Textual AnnotationZehui Li, Vallijah Subasri, Yifei Shen, Dongsheng Li 等NeurIPS 2025 · 被引用 6 次
