Reverse-Complement Equivariant Networks for DNA Sequences
Vincent Mallet, Jean-Philippe Vert
摘要
As DNA sequencing technologies keep improving in scale and cost, there is a growing need to develop machine learning models to analyze DNA sequences, e.g., to decipher regulatory signals from DNA fragments bound by a particular protein of interest. As a double helix made of two complementary strands, a DNA fragment can be sequenced as two equivalent, so-called Reverse Complement (RC) sequences of nucleotides. To take into account this inherent symmetry of the data in machine learning models can facilitate learning. In this sense, several authors have recently proposed particular RC-equivariant convolutional neural networks (CNNs). However, it remains unknown whether other RC-equivariant architectures exist, which could potentially increase the set of basic models adapted to DNA sequences for practitioners. Here, we close this gap by characterizing the set of all linear RC-equivariant layers, and show in particular that new architectures exist beyond the ones already explored. We further discuss RC-equivariant pointwise nonlinearities adapted to different architectures, as well as RC-equivariant embeddings of k-mers as an alternative to one-hot encoding of nucleotides. We show experimentally that the new architectures can outperform existing ones.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence ModelingYair Schiff, Chia-Hsiang Kao, Aaron Gokaslan, Tri Dao 等ICML 2024 · 被引用 195 次
- TrinityDNA: A Bio-Inspired Foundational Model for Efficient Long-Sequence DNA ModelingQirong Yang, Yucheng Guo, Zicheng Liu, Yujie Yang 等AAAI 2026 · 被引用 4 次
- PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNAAlice Del Vecchio, Chantriolnt-Andreas Kapourani, Abdullah M Athar, Agnieszka Dobrowolska 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper4
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
- Equivariant message passing for the prediction of tensorial properties and molecular spectraKristof Schütt, Oliver T. Unke, Michael GasteggerICML 2021 · 被引用 736 次
- On the Universality of Rotation Equivariant Point Cloud NetworksNadav Dym, Haggai MaronICLR 2021 · 被引用 22 次
相关 Paper
- Deep Squared Euclidean Approximation to the Levenshtein Distance for DNA StorageAlan J. X. Guo, Cong Liang, Qing-Hu HouICML 2022 · 被引用 5 次
- Levenshtein Distance Embedding with Poisson Regression for DNA StorageXiang Wei, Alan J. X. Guo, Sihan Sun, Mengyi Wei 等AAAI 2024 · 被引用 2 次
- Learning Structurally Stabilized Representations for Lossless DNA StorageBen Cao, Xue Li, Tiantian He, Bin Wang 等AAAI 2026
- Hyperbolic Genome EmbeddingsRaiyan R. Khan, Philippe Chlenski, Itsik Pe'erICLR 2025
- Size-Generalizable RNA Structure Evaluation by Exploring Hierarchical GeometriesZongzhao Li, Jiacheng Cen, Wenbing Huang, Taifeng Wang 等ICLR 2025
