Understanding the Limitations of Deep Models for Molecular property prediction: Insights and Solutions
Jun Xia, Lecheng Zhang, Xiao Zhu, Yue Liu, Zhangyang Gao, Bozhen Hu, Cheng Tan, Jiangbin Zheng, Siyuan Li, Stan Z. Li
摘要
Molecular Property Prediction (MPP) is a crucial task in the AI-driven Drug Discovery (AIDD) pipeline, which has recently gained considerable attention thanks to advancements in deep learning. However, recent research has revealed that deep models struggle to beat traditional non-deep ones on MPP. In this study, we benchmark 12 representative models (3 non-deep models and 9 deep models) on 15 molecule datasets. Through the most comprehensive study to date, we make the following key observations: (i) Deep models are generally unable to outperform non-deep ones; (ii) The failure of deep models on MPP cannot be solely attributed to the small size of molecular datasets; (iii) In particular, some traditional models including XGB and RF that use molecular fingerprints as inputs tend to perform better than other competitors. Furthermore, we conduct extensive empirical investigations into the unique patterns of molecule data and inductive biases of various models underlying these phenomena. These findings stimulate us to develop a simple-yet-effective feature mapping method for molecule data prior to feeding them into deep models. Empirically, deep models equipped with this mapping method can beat non-deep ones in most MoleculeNet datasets. Notably, the effectiveness is further corroborated by extensive experiments on cutting-edge dataset related to COVID-19 and activity cliff datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Representing Molecules as Random Walks Over Interpretable GrammarsMichael Sun, Minghao Guo, Weize Yuan, Veronika Thost 等ICML 2024 · 被引用 6 次
- MetaEnzyme: Meta Pan-Enzyme Learning for Task-Adaptive RedesignJiangbin Zheng, Han Zhang, Qianqing Xu, An-Ping Zeng 等ACM MM 2024 · 被引用 5 次
- Omni-Mol: Multitask Molecular Model for Any-to-any ModalitiesChengxin Hu, Hao Li, Yihe Yuan, Zezheng Song 等NeurIPS 2025 · 被引用 5 次
- TopoFormer: Topology Meets Attention for Graph LearningMd Joshem Uddin, Astrit Tola, Cuneyt Gurcan Akcora, Baris CoskunuzerICLR 2026 · 被引用 2 次
- Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware AlignmentHyuntae Park, Yeachan Kim, SangKeun LeeEMNLP 2025
它引用的顶会 Paper15
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 被引用 1,432 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
相关 Paper
- Curriculum-aware Training for Discriminating Molecular Property Prediction ModelsHansi Yang, Quanming Yao, James KwokICLR 2025
- Conformal Prediction Sets for Graph Neural NetworksSoroush H. Zargarbashi, Simone Antonelli, Aleksandar BojchevskiICML 2023 · 被引用 49 次
- Few-Shot Graph Learning for Molecular Property PredictionZhichun Guo, Chuxu Zhang, Wenhao Yu, John Herr 等WWW 2021 · 被引用 213 次
- Does GNN Pretraining Help Molecular Representation?Ruoxi Sun, Hanjun Dai, Adams Wei YuNeurIPS 2022 · 被引用 102 次
- Hierarchical Grammar-Induced Geometry for Data-Efficient Molecular Property PredictionMinghao Guo, Veronika Thost, Samuel W. Song, Adithya Balachandran 等ICML 2023
