Rigging the Foundation: Manipulating Pre-training for Advanced Membership Inference Attacks
Zihao Wang, Rui Zhu, Zhikun Zhang, Haixu Tang, XiaoFeng Wang
摘要
The significant advances in computing power have led to a surge in model complexity. Training such models today increasingly relies on transfer learning, where models are pre-trained on large datasets and later fine-tuned for different domains, allowing the knowledge in the pre-trained model to be effectively reused and customized for these specific domains. However, such a learning paradigm also opens new attack surfaces on the fine-tuned model. Particularly, a privacy risk never studied before is the threat posed by the adversary affecting the pre-training process to the downstream user's private data for fine-tuning the model: A manipulated pre-trained model can render its fine-tuned version vulnerable to privacy attacks, such as membership inference attacks (MIAs) where the presence of a given sample in the fine-tuning dataset can be determined by querying the vulnerable model. A unique challenge in understanding this privacy risk is how to amplify the membership leakage while ensuring the performance of the fine-tuned model. To address this challenge, we introduce a new technique - active robustness overfitting (ARO). This approach actively induces robustness overfitting during pre-training, which amplifies membership leakage in the downstream task without affecting its accuracy, while also maintaining the stealthiness of the attack. Our extensive evaluations across various datasets and diverse MIA scenarios demonstrate that our methods can effectively amplify membership leakage while preserving satisfactory downstream test accuracy, which contributes to a better understanding of the privacy risk introduced by transfer learning.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Privacy Auditing of Multi-Domain Graph Pre-Trained Model Under Membership Inference AttacksJiayi Luo, Qingyun Sun, Yuecen Wei, Haonan Yuan 等AAAI 2026 · 被引用 2 次
- TARP-VP: Towards Evaluation of Transferred Adversarial Robustness and Privacy on Label Mapping Visual Prompting ModelsZhen Chen, Yi Zhang, Fu Wang, Xingyu Zhao 等NeurIPS 2024 · 被引用 2 次
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang 等AAAI 2025 · 被引用 17 次
- Privacy Risks of Securing Machine Learning Models against Adversarial ExamplesLiwei Song, Reza Shokri, Prateek MittalCCS 2019 · 被引用 293 次
- When Does Data Augmentation Help With Membership Inference Attacks?Yigitcan Kaya, Tudor DumitrasICML 2021 · 被引用 82 次
