Mole-BERT: Rethinking Pre-training Graph Neural Networks for Molecules
Jun Xia, Chengshuai Zhao, Bozhen Hu, Zhangyang Gao, Cheng Tan, Yue Liu, Siyuan Li, Stan Z. Li
Abstract
Recent years have witnessed the prosperity of pre-training graph neural networks (GNNs) for molecules. Typically, atom types as node attributes are randomly masked and GNNs are then trained to predict masked types as in AttrMask , following the Masked Language Modeling (MLM) task of BERT . However, unlike MLM where the vocabulary is large, the AttrMask pre-training does not learn informative molecular representations due to small and unbalanced atom vocabulary'. To amend this problem, we propose a variant of VQ-VAE as a context-aware tokenizer to encode atom attributes into chemically meaningful discrete codes. This can enlarge the atom vocabulary size and mitigate the quantitative divergence between dominant (e.g., carbons) and rare atoms (e.g., phosphorus). With the enlarged atom vocabulary', we propose a novel node-level pre-training task, dubbed Masked Atoms Modeling (MAM), to mask some discrete codes randomly and then pre-train GNNs to predict them. MAM also mitigates another issue of AttrMask, namely the negative transfer. It can be easily combined with various pre-training tasks to improve their performance. Furthermore, we propose triplet masked contrastive learning (TMCL) for graph-level pre-training to model the heterogeneous semantic similarity between molecules for effective molecule retrieval. MAM and TMCL constitute a novel pre-training framework, Mole-BERT, which can match or outperform state-of-the-art methods in a fully data-driven manner. We release the code at magentahttps://github.com/junxia97/Mole-BERT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext abd8ff9c-1144-4f22-9223-fd3f987df0d0Cited by top-tier papers56
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
- Dink-Net: Neural Clustering on Large GraphsYue Liu, Ke Liang, Jun Xia, Sihang Zhou et al.ICML 2023 · 78 citations
- Rethinking Tokenizer and Decoder in Masked Graph Modeling for MoleculesZhiyuan Liu, Yaorui Shi, An Zhang, Enzhi Zhang et al.NeurIPS 2023 · 71 citations
- Understanding the Limitations of Deep Models for Molecular property prediction: Insights and SolutionsJun Xia, Lecheng Zhang, Xiao Zhu, Yue Liu et al.NeurIPS 2023 · 54 citations
- Learning Invariant Molecular Representation in Latent Discrete SpaceXiang Zhuang, Qiang Zhang, Keyan Ding, Yatao Bian et al.NeurIPS 2023 · 41 citations
Builds on31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- Masked Graph Modeling with Multi- View ContrastYanchen Luo, Sihang Li, Yongduo Sui, Junkang Wu et al.ICDE 2024 · 10 citations
- Dual-view Molecular Pre-trainingJinhua Zhu, Yingce Xia, Lijun Wu, Shufang Xie et al.KDD 2023 · 47 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Fragment-based Pretraining and Finetuning on Molecular GraphsKha-Dinh Luong, Ambuj K. SinghNeurIPS 2023 · 36 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
