Continuous-Discrete Convolution for Geometry-Sequence Modeling in Proteins
Hehe Fan, Zhangyang Wang, Yi Yang, Mohan S. Kankanhalli
摘要
The structure of proteins involves 3D geometry of amino acid coordinates and 1D sequence of peptide chains. The 3D structure exhibits irregularity because amino acids are distributed unevenly in Euclidean space and their coordinates are continuous variables. In contrast, the 1D structure is regular because amino acids are arranged uniformly in the chains and their sequential positions (orders) are discrete variables. Moreover, geometric coordinates and sequential orders are in two types of spaces and their units of length are incompatible. These inconsistencies make it challenging to capture the 3D and 1D structures while avoiding the impact of sequence and geometry modeling on each other. This paper proposes a Continuous-Discrete Convolution (CDConv) that uses irregular and regular approaches to model the geometry and sequence structures, respectively. Specifically, CDConv employs independent learnable weights for different regular sequential displacements but directly encodes geometric displacements due to their irregularity. In this way, CDConv significantly improves protein modeling by reducing the impact of geometric irregularity on sequence modeling. Extensive experiments on a range of tasks, including protein fold classification, enzyme reaction classification, gene ontology term prediction and enzyme commission number prediction, demonstrate the effectiveness of the proposed CDConv.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper20
- FoldToken: Learning Protein Language via Vector Quantization and BeyondZhangyang Gao, Cheng Tan, Jue Wang, Yufei Huang 等AAAI 2025 · 被引用 29 次
- Evaluating Representation Learning on the Protein Structure UniverseArian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang 等ICLR 2024 · 被引用 26 次
- Clustering for Protein Representation LearningRuijie Quan, Wenguan Wang, Fan Ma, Hehe Fan 等CVPR 2024 · 被引用 10 次
- ProtGO: Function-Guided Protein Modeling for Unified Representation LearningBozhen Hu, Cheng Tan, Yongjie Xu, Zhangyang Gao 等NeurIPS 2024 · 被引用 10 次
- Pre-Training Protein Bi-level Representation Through Span Mask Strategy On 3D Protein ChainsJiale Zhao, Wanru Zhuang, Jia Song, Yaqi Li 等ICML 2024 · 被引用 9 次
相关 Paper
- Geometric Graph Representation Learning on Protein Structure PredictionTian Xia, Wei-Shinn KuKDD 2021 · 被引用 28 次
- Protein Representation Learning by Geometric Structure PretrainingZuobai Zhang, Minghao Xu, Arian Rokkum Jamasb, Vijil Chenthamarakshan 等ICLR 2023 · 被引用 40 次
- Fast End-to-End Learning on Protein SurfacesFreyr Sverrisson, Jean Feydy, Bruno E. Correia, Michael M. BronsteinCVPR 2021
- Intrinsic-Extrinsic Convolution and Pooling for Learning on 3D Protein StructuresPedro Hermosilla, Marco Schäfer, Matej Lang, Gloria Fackelmann 等ICLR 2021 · 被引用 119 次
- Generative Flows on Discrete State-Spaces: Enabling Multimodal Flows with Applications to Protein Co-DesignAndrew Campbell, Jason Yim, Regina Barzilay, Tom Rainforth 等ICML 2024 · 被引用 283 次
