Motif-based Graph Self-Supervised Learning for Molecular Property Prediction
Zaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu, Chee-Kong Lee
Abstract
Predicting molecular properties with data-driven methods has drawn much attention in recent years. Particularly, Graph Neural Networks (GNNs) have demonstrated remarkable success in various molecular generation and prediction tasks. In cases where labeled data is scarce, GNNs can be pre-trained on unlabeled molecular data to first learn the general semantic and structural information before being fine-tuned for specific tasks. However, most existing self-supervised pre-training frameworks for GNNs only focus on node-level or graph-level tasks. These approaches cannot capture the rich information in subgraphs or graph motifs. For example, functional groups (frequently-occurred subgraphs in molecular graphs) often carry indicative information about the molecular properties. To bridge this gap, we propose Motif-based Graph Self-supervised Learning (MGSSL) by introducing a novel self-supervised motif generation framework for GNNs. First, for motif extraction from molecular graphs, we design a molecule fragmentation method that leverages a retrosynthesis-based algorithm BRICS and additional rules for controlling the size of motif vocabulary. Second, we design a general motif-based generative pre-training framework in which GNNs are asked to make topological and label predictions. This generative framework can be implemented in two different ways, i.e., breadth-first or depth-first. Finally, to take the multi-scale information in molecular graphs into consideration, we introduce a multi-level self-supervised pre-training. Extensive experiments on various downstream benchmark tasks show that our methods outperform all state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers78
- ProtGNN: Towards Self-Explaining Graph Neural NetworksZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.AAAI 2022 · 173 citations
- Universal Prompt Tuning for Graph Neural NetworksTaoran Fang, Yunchao Zhang, Yang Yang, Chunping Wang et al.NeurIPS 2023 · 166 citations
- Hierarchical Graph Transformer with Adaptive Node SamplingZaixi Zhang, Qi Liu, Qingyong Hu, Chee-Kong LeeNeurIPS 2022 · 145 citations
- Mole-BERT: Rethinking Pre-training Graph Neural Networks for MoleculesJun Xia, Chengshuai Zhao, Bozhen Hu, Zhangyang Gao et al.ICLR 2023 · 119 citations
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- Directional Message Passing for Molecular GraphsJohannes Klicpera, Janek Groß, Stephan GünnemannICLR 2020 · 1,079 citations
Related papers
- Fragment-based Pretraining and Finetuning on Molecular GraphsKha-Dinh Luong, Ambuj K. SinghNeurIPS 2023 · 36 citations
- Pre-Training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information BottleneckVan Thuy Hoang, O-Joun LeeAAAI 2025 · 19 citations
- KPGT: Knowledge-Guided Pre-training of Graph Transformer for Molecular Property PredictionHan Li, Dan Zhao, Jianyang ZengKDD 2022 · 55 citations
- Few-Shot Graph Learning for Molecular Property PredictionZhichun Guo, Chuxu Zhang, Wenhao Yu, John Herr et al.WWW 2021 · 213 citations
- MAGE: Model-Level Graph Neural Networks Explanations via Motif-based Graph GenerationZhaoning Yu, Hongyang GaoICLR 2025
