CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity Prediction
Rong Han, Xiaohong Liu, Tong Pan, Jing Xu, Xiaoyu Wang, Wuyang Lan, Zhenyu Li, Zixuan Wang, Jiangning Song, Guangyu Wang, Ting Chen
摘要
Accurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. Co-PRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MSA TransformerRoshan Rao, Jason Liu, Robert Verkuil, Joshua Meier 等ICML 2021 · 被引用 686 次
- Learning from Protein Structure with Geometric Vector PerceptronsBowen Jing, Stephan Eismann, Patricia Suriana, Raphael John Lamarre Townshend 等ICLR 2021 · 被引用 627 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
相关 Paper
- CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesLinglin Jing, Sheng Xu, Yifan Wang, Yuzhe Zhou 等AAAI 2024 · 被引用 9 次
- Large Language and Protein Assistant for Protein-Protein Interactions PredictionPeng Zhou, Pengsen Ma, Jianmin Wang, Xibao Cai 等ACL 2025 · 被引用 2 次
- RBPtool: A Deep Language Model Framework for Multi-Resolution RBP-RNA Binding Prediction and RNA Molecule DesignJiyue Jiang, Yitao Xu, Zikang Wang, Yihan Ye 等EMNLP 2025
- ProtCLIP: Function-Informed Protein Multi-Modal LearningHanjing Zhou, Mingze Yin, Wei Wu, Mingyang Li 等AAAI 2025 · 被引用 11 次
- Towards All-Atom Foundation Models for Biomolecular Binding Affinity PredictionLiang Shi, Zuobai Zhang, Huiyu Cai, Santiago Miret 等ICLR 2026
