Molecular Data Programming: Towards Molecule Pseudo-labeling with Systematic Weak Supervision
Xin Juan, Kaixiong Zhou, Ninghao Liu, Tianlong Chen, Xin Wang
Abstract
The premise for the great advancement of molecular machine learning is dependent on a considerable amount of labeled data. In many real-world scenarios, the labeled molecules are limited in quantity or laborious to derive. Recent pseudo-labeling methods are usually designed based on a single domain knowledge, thereby failing to understand the comprehensive molecular configurations and limiting their adaptability to generalize across diverse biochemical context. To this end, we introduce an innovative paradigm for dealing with the molecule pseudo-labeling, named as Molecular Data Programming (MDP). In particular, we adopt systematic supervision sources via crafting multiple graph labeling functions, which covers various molecular structural knowledge of graph kernels, molecular fingerprints, and topological features. Each of them creates an uncertain and biased labels for the unlabeled molecules. To address the decision conflicts among the diverse pseudo-labels, we design a label synchronizer to differentiably model confidences and correlations between the labeling functions, which yields probabilistic molecular labels to adapt for specific applications. These probabilistic molecular labels are used to train a molecular classifier for improving its generalization capability. On eight benchmark datasets, we empirically demonstrate the effectiveness of MDP on the weakly supervised molecule classification tasks, achieving an average improvement of 9.5%. The code is in: https://github.com/xinjuan1/MDP/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80e12bfe-1779-4ba7-af4e-f2c29e1371d9Cited by top-tier papers1
Ask how each one uses itBuilds on12
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-supervised Graph-level Representation Learning with Local and Global StructureMinghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo et al.ICML 2021 · 248 citations
- Directional Graph NetworksDominique Beaini, Saro Passaro, Vincent Létourneau, William L. Hamilton et al.ICML 2021 · 216 citations
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang et al.KDD 2022 · 141 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
Related papers
- DP-SSL: Towards Robust Semi-supervised Learning with A Few Labeled SamplesYi Xu, Jiandong Ding, Lu Zhang, Shuigeng ZhouNeurIPS 2021 · 34 citations
- Instructor-inspired Machine Learning for Robust Molecular Property PredictionFang Wu, Shuting Jin, Siyuan Li, Stan Z. LiNeurIPS 2024 · 14 citations
- KGOT: Unified Knowledge Graph and Optimal Transport Pseudo-Labeling for Molecule-Protein Interaction PredictionJiayu Qin, Zhengquan Luo, Guy Tadmor, Changyou Chen et al.ICLR 2026 · 2 citations
- Robust Data Programming with Precision-guided Labeling FunctionsOishik Chatterjee, Ganesh Ramakrishnan, Sunita SarawagiAAAI 2020 · 20 citations
- Witan: Unsupervised Labelling Function Generation for Assisted Data ProgrammingBenjamin Denham, Edmund M.-K. Lai, Roopak Sinha, Muhammad Asif NaeemVLDB 2022 · 12 citations
