MORE: Molecule Pretraining with Multi-Level Pretext Task
Yeongyeong Son, Dasom Noh, Gyoungyoung Heo, Gyoung Jin Park, Sunyoung Kwon
Abstract
Foundation models, serving as pretrained fundamental bases for a variety of downstream tasks, try to learn versatile, rich, and generalizable representations that can be quickly adopted through fine-tuning or even in a zero-shot manner for specific applications. Foundation models for molecular representation are no exception. Various pretext tasks have been proposed for pretraining molecular representations, but these approaches have focused on only single or partial properties. Molecules are complicated and require different perspectives depending on purposes: insights from local- or global-level, 2D-topology or 3D-spatial arrangement, and low- or high-level semantics. We propose Multi-level mOlecule gRaph prE-train (MORE) to consider these multiple aspects of molecules simultaneously. Experimental results demonstrate that our proposed method effectively learns comprehensive representations by showing outstanding performance in both linear probing and full fine-tuning. Notably, in quantification experiments of forgetting the pretrained models, MORE consistently exhibits minimal and stable parameter changes with the smallest performance gap, whereas other methods show substantial and inconsistent fluctuations with larger gaps. The effectiveness of individual pretext tasks varies depending on the problems being solved, which again highlights the need for a multi-level perspective. Scalability experiments reveal steady improvements of MORE as the dataset size increases, suggesting potential gains with larger datasets as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76cccca3-2b62-4967-8b30-4e6f08c05e2eCited by top-tier papers1
Ask how each one uses itBuilds on9
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
- Pre-training Molecular Graph Representation with 3D GeometryShengchao Liu, Hanchen Wang, Weiyang Liu, Joan Lasenby et al.ICLR 2022 · 440 citations
- Motif-based Graph Self-Supervised Learning for Molecular Property PredictionZaixi Zhang, Qi Liu, Hao Wang, Chengqiang Lu et al.NeurIPS 2021 · 385 citations
- Self-supervised Graph-level Representation Learning with Local and Global StructureMinghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo et al.ICML 2021 · 248 citations
Related papers
- UniCorn: A Unified Contrastive Learning Approach for Multi-view Molecular Representation LearningShikun Feng, Yuyan Ni, Minghao Li, Yanwen Huang et al.ICML 2024 · 22 citations
- 3D Infomax improves GNNs for Molecular Property PredictionHannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou et al.ICML 2022 · 269 citations
- MIPT: Multilevel Informed Prompt Tuning for Robust Molecular Property PredictionYeyun Chen, Jiangming ShiICML 2025
- Advancing Molecular Graph-Text Pre-training via Fine-grained AlignmentYibo Li, Yuan Fang, Mengmei Zhang, Chuan ShiKDD 2025
- Energy-Motivated Equivariant Pretraining for 3D Molecular GraphsRui Jiao, Jiaqi Han, Wenbing Huang, Yu Rong et al.AAAI 2023 · 64 citations
