Extracting Training Data from Molecular Pre-trained Models
Renhong Huang, Jiarong Xu, Zhiming Yang, Xiang Si, Xin Jiang, Hanyang Yuan, Chunping Wang, Yang Yang
Abstract
Graph Neural Networks (GNNs) have significantly advanced the field of drug discovery, enhancing the speed and efficiency of molecular identification. However, training these GNNs demands vast amounts of molecular data, which has spurred the emergence of collaborative model-sharing initiatives. These initiatives facilitate the sharing of molecular pre-trained models among organizations without exposing proprietary training data. Despite the benefits, these molecular pre-trained models may still pose privacy risks. For example, malicious adversaries could perform data extraction attack to recover private training data, thereby threatening commercial secrets and collaborative trust. This work, for the first time, explores the risks of extracting private training molecular data from molecular pre-trained models. This task is nontrivial as the molecular pre-trained models are non-generative and exhibit a diversity of model architectures, which differs significantly from language and image models. To address these issues, we introduce a molecule generation approach and propose a novel, model-independent scoring function for selecting promising molecules. To efficiently reduce the search space of potential molecules, we further introduce a Molecule Extraction Policy Network for molecule extraction. Our experiments demonstrate that even with only query access to molecular pre-trained models, there is a considerable risk of extracting training data, challenging the assumption that model sharing alone provides adequate protection against data extraction attacks. Our codes are publicly available at: https://github.com/ Molextract/Data-Extraction-from-Molecular-Pre-trained-Model .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec581b98-9b4a-4821-9b42-d50e9a5b5a81Cited by top-tier papers2
- AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMsLara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas CavigelliNeurIPS 2025 · 3 citations
- KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural NetworksTaoran Fang, Tianhong Gao, Chunping Wang, Yihao Shang et al.ICLR 2025 · 1 citation
Builds on24
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie et al.NeurIPS 2020 · 1,113 citations
Related papers
- Unveiling the Secrets without Data: Can Graph Neural Networks Be Exploited through Data-Free Model Extraction Attacks?Yuanxin Zhuang, Chuan Shi, Mengmei Zhang, Jinghui Chen et al.USENIX Security 2024 · 11 citations
- Can Graph Neural Networks Expose Training Data Properties? An Efficient Risk Assessment ApproachHanyang Yuan, Jiarong Xu, Renhong Huang, Mingli Song et al.NeurIPS 2024 · 3 citations
- Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing of Graph ModelsBin Ma, Yuyuan Feng, Minhua Lin, Enyan DaiKDD 2026 · 1 citation
- Does GNN Pretraining Help Molecular Representation?Ruoxi Sun, Hanjun Dai, Adams Wei YuNeurIPS 2022 · 102 citations
- GrOVe: Ownership Verification of Graph Neural Networks using EmbeddingsAsim Waheed, Vasisht Duddu, N. AsokanS&P 2024 · 19 citations
