Extracting Training Data from Molecular Pre-trained Models
Renhong Huang, Jiarong Xu, Zhiming Yang, Xiang Si, Xin Jiang, Hanyang Yuan, Chunping Wang, Yang Yang
摘要
Graph Neural Networks (GNNs) have significantly advanced the field of drug discovery, enhancing the speed and efficiency of molecular identification. However, training these GNNs demands vast amounts of molecular data, which has spurred the emergence of collaborative model-sharing initiatives. These initiatives facilitate the sharing of molecular pre-trained models among organizations without exposing proprietary training data. Despite the benefits, these molecular pre-trained models may still pose privacy risks. For example, malicious adversaries could perform data extraction attack to recover private training data, thereby threatening commercial secrets and collaborative trust. This work, for the first time, explores the risks of extracting private training molecular data from molecular pre-trained models. This task is nontrivial as the molecular pre-trained models are non-generative and exhibit a diversity of model architectures, which differs significantly from language and image models. To address these issues, we introduce a molecule generation approach and propose a novel, model-independent scoring function for selecting promising molecules. To efficiently reduce the search space of potential molecules, we further introduce a Molecule Extraction Policy Network for molecule extraction. Our experiments demonstrate that even with only query access to molecular pre-trained models, there is a considerable risk of extracting training data, challenging the assumption that model sharing alone provides adequate protection against data extraction attacks. Our codes are publicly available at: https://github.com/ Molextract/Data-Extraction-from-Molecular-Pre-trained-Model .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AC-LoRA: (Almost) Training-Free Access Control Aware Multi-Modal LLMsLara Magdalena Lazier, Aritra Dhar, Vasilije Stambolic, Lukas CavigelliNeurIPS 2025 · 被引用 3 次
- KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural NetworksTaoran Fang, Tianhong Gao, Chunping Wang, Yihao Shang 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper24
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- Self-Supervised Graph Transformer on Large-Scale Molecular DataYu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie 等NeurIPS 2020 · 被引用 1,113 次
相关 Paper
- Unveiling the Secrets without Data: Can Graph Neural Networks Be Exploited through Data-Free Model Extraction Attacks?Yuanxin Zhuang, Chuan Shi, Mengmei Zhang, Jinghui Chen 等USENIX Security 2024 · 被引用 11 次
- Can Graph Neural Networks Expose Training Data Properties? An Efficient Risk Assessment ApproachHanyang Yuan, Jiarong Xu, Renhong Huang, Mingli Song 等NeurIPS 2024 · 被引用 3 次
- Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing of Graph ModelsBin Ma, Yuyuan Feng, Minhua Lin, Enyan DaiKDD 2026 · 被引用 1 次
- Does GNN Pretraining Help Molecular Representation?Ruoxi Sun, Hanjun Dai, Adams Wei YuNeurIPS 2022 · 被引用 102 次
- GrOVe: Ownership Verification of Graph Neural Networks using EmbeddingsAsim Waheed, Vasisht Duddu, N. AsokanS&P 2024 · 被引用 19 次
