Fragment-based Pretraining and Finetuning on Molecular Graphs
Kha-Dinh Luong, Ambuj K. Singh
Abstract
Property prediction on molecular graphs is an important application of Graph Neural Networks. Recently, unlabeled molecular data has become abundant, which facilitates the rapid development of self-supervised learning for GNNs in the chemical domain. In this work, we propose pretraining GNNs at the fragment level, a promising middle ground to overcome the limitations of node-level and graph-level pretraining. Borrowing techniques from recent work on principal subgraph mining, we obtain a compact vocabulary of prevalent fragments from a large pretraining dataset. From the extracted vocabulary, we introduce several fragment-based contrastive and predictive pretraining tasks. The contrastive learning task jointly pretrains two different GNNs: one on molecular graphs and the other on fragment graphs, which represents higher-order connectivity within molecules. By enforcing consistency between the fragment embedding and the aggregated embedding of the corresponding atoms from the molecular graphs, we ensure that the embeddings capture structural information at multiple resolutions. The structural information of fragment graphs is further exploited to extract auxiliary labels for graph-level predictive pretraining. We employ both the pretrained molecular-based and fragment-based GNNs for downstream prediction, thus utilizing the fragment information during finetuning. Our graph fragment-based pretraining (GraphFP) advances the performances on 5 out of 8 common molecular benchmarks and improves the performances on long-range biological benchmarks by at least 11.5%. Code is available at: https://github.com/lvkd84/GraphFP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Pre-Training Graph Neural Networks on Molecules by Using Subgraph-Conditioned Graph Information BottleneckVan Thuy Hoang, O-Joun LeeAAAI 2025 · 19 citations
- Hierarchical Graph Tokenization for Molecule-Language AlignmentYongqiang Chen, Quanming Yao, Juzheng Zhang, James Cheng et al.ICML 2025 · 2 citations
- Learning the Neighborhood: Contrast-Free Multimodal Self-Supervised Molecular Graph PretrainingBoshra Ariguib, Mathias Niepert, Andrei ManolacheICML 2026
Builds on20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- Motif-based Graph Self-Supervised Learning for Molecular Property PredictionZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.NeurIPS 2021 · 385 citations
- GCC: Graph Contrastive Coding for Graph Neural Network Pre-TrainingJiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang et al.KDD 2020 · 755 citations
- Learning to Pre-train Graph Neural NetworksYuanfu Lu, Xunqiang Jiang, Yuan Fang, Chuan ShiAAAI 2021 · 158 citations
- Bi-level Contrastive Learning for Knowledge-Enhanced Molecule RepresentationsPengcheng Jiang, Cao Xiao, Tianfan Fu, Parminder Bhatia et al.AAAI 2025 · 7 citations
- Does GNN Pretraining Help Molecular Representation?Ruoxi Sun, Hanjun Dai, Adams Wei YuNeurIPS 2022 · 102 citations
