Privacy Backdoors: Stealing Data with Corrupted Pretrained Models
Shanglun Feng, Florian Tramèr
摘要
Practitioners commonly download pretrained machine learning models from open repositories and finetune them to fit specific applications. We show that this practice introduces a new risk of privacy backdoors. By tampering with a pretrained model's weights, an attacker can fully compromise the privacy of the finetuning data. We show how to build privacy backdoors for a variety of models, including transformers, which enable an attacker to reconstruct individual finetuning samples, with a guaranteed success! We further show that backdoored models allow for tight privacy attacks on models trained with differential privacy (DP). The common optimistic practice of training DP models with loose privacy guarantees is thus insecure if the model is not trusted. Overall, our work highlights a crucial and overlooked supply chain attack on machine learning privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsYuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping 等NeurIPS 2024 · 被引用 39 次
- Nearly Tight Black-Box Auditing of Differentially Private Machine LearningMeenatchi Sundaram Muthu Selva Annamalai, Emiliano De CristofaroNeurIPS 2024 · 被引用 32 次
- Evaluations of Machine Learning Privacy Defenses are MisleadingMichael Aerni, Jie Zhang, Florian TramèrCCS 2024 · 被引用 12 次
- BEEAR: Embedding-based Adversarial Removal of Safety Backdoors in Instruction-tuned Language ModelsYi Zeng, Weiyu Sun, Tran Ngoc Huynh, Dawn Song 等EMNLP 2024 · 被引用 10 次
- PreCurious: How Innocent Pre-Trained Language Models Turn into Privacy TrapsRuixuan Liu, Tianhao Wang, Yang Cao, Li XiongCCS 2024 · 被引用 9 次
它引用的顶会 Paper21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
相关 Paper
- Handcrafted Backdoors in Deep Neural NetworksSanghyun Hong, Nicholas Carlini, Alexey KurakinNeurIPS 2022 · 被引用 105 次
- Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding IndistinguishabilityHao Wang, Shangwei Guo, Jialing He, Hangcheng Liu 等WWW 2025 · 被引用 10 次
- Dormant Backdoor: Weaponizing Model Finetuning for Feasible Backdoor Attacks Against Pretrained ModelsRuitao Li, Jiakai Wang, Hairong Chen, Huihu Ding 等AAAI 2026
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy LeakageMd. Rafi Ur Rashid, Jing Liu, Toshiaki Koike-Akino, Ye Wang 等AAAI 2025 · 被引用 17 次
