GRAM-DTI: Adaptive Multimodal Representation Learning for Drug–Target Interaction Prediction
Feng Jiang, Amina Mollaysa, Hehuan Ma, Yuzhi Guo, Tommaso Mansi, Junzhou Huang, Mangal Prakash, Rui Liao
Abstract
Drug target interaction (DTI) prediction is a cornerstone of computational drug discovery, enabling rational design, repurposing, and mechanistic insights. While deep learning has advanced DTI modeling, existing approaches primarily rely on SMILES–protein pairs and fail to exploit the rich multimodal information available for small molecules and proteins. Inspired by recent successes in multimodal molecular property prediction, we introduce GRAM-DTI, a pre-training framework that integrates multimodal small molecule and protein inputs into a unified representation. GRAM-DTI extends volume-based contrastive learning to four modalities, capturing higher-order semantic alignment beyond conventional pairwise approaches. To handle modality informativeness, we propose adaptive modality dropout, dynamically regulating each modality’s contribution during pretraining. Additionally, IC50 activity measurements, when available, are incorporated as weak supervision to ground representations in biologically meaningful interaction strengths. Experiments on four publicly available datasets demonstrate that GRAM-DTI consistently outperforms state-of-the-art baselines. Our results highlight the benefits of higher-order multimodal alignment, adaptive modality utilization, and auxiliary supervision for robust and generalizable DTI prediction. Our code is available at https://github.com/uta-smile/GRAM-DTI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b98a9e1-4a61-4090-b5c3-830a3c7331bbCited by top-tier papers1
Ask how each one uses itBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- EquiBind: Geometric Deep Learning for Drug Binding Structure PredictionHannes Stärk, Octavian Ganea, Lagnajit Pattanaik, Regina Barzilay et al.ICML 2022 · 360 citations
- Translation between Molecules and Natural LanguageCarl Edwards, Tuan Manh Lai, Kevin Ros, Garrett Honke et al.EMNLP 2022 · 112 citations
- MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal AdapterZhiyuan Liu, Sihang Li, Yanchen Luo, Hao Fei et al.EMNLP 2023 · 34 citations
- Learnable Irrelevant Modality Dropout for Multimodal Action Recognition on Modality-Specific Annotated VideosSaghir Alfasly, Jian Lu, Chen Xu, Yuru ZouCVPR 2022 · 28 citations
Related papers
- Generalizable Drug-Target Interaction Prediction via ESM-2 Representations and Progressive Contrastive Curriculum LearningQianyang Wu, Jingwei Lv, Zilong Zhang, Feifei CuiAAAI 2026
- R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance ExplorationYang Hua, Tianyang Xu, Xiaoning Song, Zhenhua Feng et al.AAAI 2025 · 2 citations
- Multi-view Graph Contrastive Representation Learning for Drug-Drug Interaction PredictionYingheng Wang, Yaosen Min, Xin Chen, Ji WuWWW 2021 · 186 citations
- TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local CorrespondenceFeng Jiang, Mangal Prakash, Hehuan Ma, Jianyuan Deng et al.NeurIPS 2025 · 5 citations
- PSC-CPI: Multi-Scale Protein Sequence-Structure Contrasting for Efficient and Generalizable Compound-Protein Interaction PredictionLirong Wu, Yufei Huang, Cheng Tan, Zhangyang Gao et al.AAAI 2024 · 20 citations
