Sequence-Free for Compound Protein Interaction Prediction
Hongzhi Zhang, Jiameng Chen, Kun Li, Yida Xiong, Xiantao Cai, Wenbin Hu, Jia Wu
Abstract
The prediction of compound–protein interactions (CPIs) is crucial for drug discovery. Most existing CPI prediction models rely on protein sequence information as input. However, in early-stage drug development, particularly in phenotype-driven studies or compound-response analyses, proteins are often annotated only with functional labels, and their sequences remain undetermined. Consequently, current methods are inapplicable in such scenarios. Furthermore, our experiments find that even when large-scale perturbations were applied to protein sequences, the predictive performance of the existing models did not show a significant decline. It indicates that the high investment in sequencing may not bring corresponding returns. To address the above issues, we propose an inexpensive, protein-sequencing-free framework BioText-CPI, based on the Biomedical Textual description of protein for CPI prediction. Firstly, during the pre-training stage of the model, we use contrastive learning to align protein texts and sequence modalities. Subsequently, we add biological text descriptions of proteins to the existing public CPI dataset to construct a new CPI dataset. Finally, in the CPI prediction stage, the sequence and biomedical text descriptions of proteins can be used as the input for CPI prediction either separately or simultaneously to meet the application requirements of different scenarios. The experiments demonstrate that BioText-CPI achieves comparable effects to the traditional methods when only the biomedical description of protein is input. Moreover, when the two modalities of protein information are input simultaneously, BioText-CPI achieves state-of-the-art performance across multiple scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Language models enable zero-shot prediction of the effects of mutations on protein functionJoshua Meier, Roshan Rao, Robert Verkuil, Jason Liu et al.NeurIPS 2021 · 969 citations
- TANKBind: Trigonometry-Aware Neural NetworKs for Drug-Protein Binding Structure PredictionWei Lu, Qifeng Wu, Jixian Zhang, Jiahua Rao et al.NeurIPS 2022 · 254 citations
- ProtST: Multi-Modality Learning of Protein Sequences and Biomedical TextsMinghao Xu, Xinyu Yuan, Santiago Miret, Jian TangICML 2023 · 147 citations
- Improving the Gating Mechanism of Recurrent Neural NetworksAlbert Gu, Çaglar Gülçehre, Thomas Paine, Matt Hoffman et al.ICML 2020 · 111 citations
- Knowledge Distillation Improves Graph Structure Augmentation for Graph Neural NetworksLirong Wu, Haitao Lin, Yufei Huang, Stan Z. LiNeurIPS 2022 · 60 citations
Related papers
- PSC-CPI: Multi-Scale Protein Sequence-Structure Contrasting for Efficient and Generalizable Compound-Protein Interaction PredictionLirong Wu, Yufei Huang, Cheng Tan, Zhangyang Gao et al.AAAI 2024 · 20 citations
- S²Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual ScreeningBowei He, Bowen Gao, Yankai Chen, Yanyan Lan et al.AAAI 2026 · 1 citation
- Prot2Text-V2: Protein Function Prediction with Multimodal Contrastive AlignmentXiao Fei, Michail Chatzianastasis, Sarah Almeida Carneiro, Hadi Abdine et al.NeurIPS 2025 · 12 citations
- FuseMine: Robust Multi-Modal Compound-Protein Interaction Prediction via Differential Attention Feature MiningJunlin Xu, Zhuang Zhang, Zhenghang Gong, Jincan Li et al.AAAI 2026
- GRAM-DTI: Adaptive Multimodal Representation Learning for Drug–Target Interaction PredictionFeng Jiang, Amina Mollaysa, Hehuan Ma, Yuzhi Guo et al.ICLR 2026
