Pre-training via Denoising for Molecular Property Prediction
Sheheryar Zaidi, Michael Schaarschmidt, James Martens, Hyunjik Kim, Yee Whye Teh, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Razvan Pascanu, Jonathan Godwin
Abstract
Many important problems involving molecular property prediction from 3D structures have limited data, posing a generalization challenge for neural networks. In this paper, we describe a pre-training technique based on denoising that achieves a new state-of-the-art in molecular property prediction by utilizing large datasets of 3D molecular structures at equilibrium to learn meaningful representations for downstream tasks. Relying on the well-known link between denoising autoencoders and score-matching, we show that the denoising objective corresponds to learning a molecular force field -- arising from approximating the Boltzmann distribution with a mixture of Gaussians -- directly from equilibrium structures. Our experiments demonstrate that using this pre-training objective significantly improves performance on multiple benchmarks, achieving a new state-of-the-art on the majority of targets in the widely used QM9 dataset. Our analysis then provides practical insights into the effects of different factors -- dataset sizes, model size and architecture, and the choice of upstream and downstream datasets -- on pre-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c0c6b8d-d28a-4f06-bda1-b1f7adaaa2faCited by top-tier papers58
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsYi-Lun Liao, Brandon M. Wood, Abhishek Das, Tess E. SmidtICLR 2024 · 311 citations
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
- From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property PredictionNima Shoghi, Adeesh Kolluru, John R. Kitchin, Zachary W. Ulissi et al.ICLR 2024 · 63 citations
- Swallowing the Bitter Pill: Simplified Scalable Conformer GenerationYuyang Wang, Ahmed A. A. Elhag, Navdeep Jaitly, Joshua M. Susskind et al.ICML 2024 · 55 citations
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 52 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
Related papers
- DenoiseVAE: Learning Molecule-Adaptive Noise Distributions for Denoising-based 3D Molecular Pre-trainingYurou Liu, Jiahao Chen, Rui Jiao, Jiangmeng Li et al.ICLR 2025
- MOES-Pred: Molecular Structural Representation Learning by Adaptive Energy-Sentinel Vibration for Generalized Property PredictionZHIRAN HOU, TINGHUAI MA, Huan Rong, Li Jia et al.ICML 2026
- May the Force be with You: Unified Force-Centric Pre-Training for 3D Molecular ConformationsRui Feng, Qi Zhu, Huan Tran, Binghong Chen et al.NeurIPS 2023 · 16 citations
- Energy-Motivated Equivariant Pretraining for 3D Molecular GraphsRui Jiao, Jiaqi Han, Wenbing Huang, Yu Rong et al.AAAI 2023 · 64 citations
- Coordinate Denoising for Non‑Equilibrium Molecular Representation LearningQianwei Tang, Baile Xu, Jian Zhao, Furao ShenCVPR 2026
