Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task Learning
Yuxuan Ren, Dihan Zheng, Chang Liu, Peiran Jin, Yu Shi, Lin Huang, Jiyan He, Shengjie Luo, Tao Qin, Tie-Yan Liu
Abstract
In recent years, machine learning has demonstrated impressive capability in handling molecular science tasks. To support various molecular properties at scale, machine learning models are trained in the multi-task learning paradigm. Nevertheless, data of different molecular properties are often not aligned: some quantities, e.g. equilibrium structure, demand more cost to compute than others, e.g. energy, so their data are often generated by cheaper computational methods at the cost of lower accuracy, which cannot be directly overcome through multi-task learning. Moreover, it is not straightforward to leverage abundant data of other tasks to benefit a particular task. To handle such data heterogeneity challenges, we exploit the specialty of molecular tasks that there are physical laws connecting them, and design consistency training approaches that allow different tasks to exchange information directly so as to improve one another. Particularly, we demonstrate that the more accurate energy data can improve the accuracy of structure prediction. We also find that consistency training can directly leverage force and off-equilibrium structure data to improve structure prediction, demonstrating a broad capability for integrating heterogeneous data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d9d62d8-f3ac-4a7c-b054-a4b6ebbcb3bbCited by top-tier papers1
Ask how each one uses itBuilds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force FieldsIlyes Batatia, Dávid Péter Kovács, Gregor N. C. Simm, Christoph Ortner et al.NeurIPS 2022 · 1,448 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Self-Consistency Training for Density-Functional-Theory Hamiltonian PredictionHe Zhang, Chang Liu, Zun Wang, Xinran Wei et al.ICML 2024 · 13 citations
- Semi-Supervised Learning for Molecular Graphs via Ensemble ConsensusRasmus Tirsgaard, Laurits Fredsgaard, Marisa Wodrich, Mikkel Jordahn et al.ICML 2026 · 1 citation
- Automatic Auxiliary Task Selection and Adaptive Weighting Boost Molecular Property PredictionZhiqiang Zhong, Davide MottinNeurIPS 2025 · 3 citations
- HeMeNet: Heterogeneous Multichannel Equivariant Network for Protein Multi-task LearningRong Han, Wenbing Huang, Lingxiao Luo, Xinyan Han et al.AAAI 2025
- Few-Shot Graph Learning for Molecular Property PredictionZhichun Guo, Chuxu Zhang, Wenhao Yu, John Herr et al.WWW 2021 · 213 citations
