Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN Models
Shawn Shan, Wenxin Ding, Emily Wenger, Haitao Zheng, Ben Y. Zhao
摘要
Server breaches are an unfortunate reality on today's Internet. In the context of deep neural network (DNN) models, they are particularly harmful, because a leaked model gives an attacker "whitebox" access to generate adversarial examples, a threat model that has no practical robust defenses. For practitioners who have invested years and millions into proprietary DNNs, e.g. medical imaging, this seems like an inevitable disaster looming on the horizon. In this paper, we consider the problem of post-breach recovery for DNN models. We propose Neo, a new system that creates new versions of leaked models, alongside an inference time filter that detects and removes adversarial examples generated on previously leaked models. The classification surfaces of different model versions are slightly offset (by introducing hidden distributions), and Neo detects the overfitting of attacks to the leaked model used in its generation. We show that across a variety of tasks and attack methods, Neo is able to filter out attacks from leaked models with very high accuracy, and provides strong protection (7-10 recoveries) against attackers who repeatedly breach the server. Neo performs well against a variety of strong adaptive attacks, dropping slightly in # of breaches recoverable, and demonstrates potential as a complement to DNN defenses in the wild.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren 等CCS 2023 · 被引用 19 次
- FaceObfuscator: Defending Deep Learning-based Privacy Attacks with Gradient Descent-resistant Features in Face RecognitionShuaifan Jin, He Wang, Zhibo Wang, Feng Xiao 等USENIX Security 2024 · 被引用 9 次
- Fight Fire with Fire: Combating Adversarial Patch Attacks using Pattern-randomized Defensive PatchesJianan Feng, Jiachun Li, Changqing Miao, Jianjun Huang 等S&P 2025
- Glaze: Protecting Artists from Style Mimicry by Text-to-Image ModelsShawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng 等USENIX Security 2023
- Enhance Stealthiness and Transferability of Adversarial Attacks with Class Activation Mapping Ensemble AttackHui Xia, Rui Zhang, Zi Kang, Shuliang Jiang 等NDSS 2024
它引用的顶会 Paper23
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
相关 Paper
- Eliminating Adversarial Noise via Information Discard and Robust Representation RestorationDawei Zhou, Yukun Chen, Nannan Wang, Decheng Liu 等ICML 2023 · 被引用 10 次
- Defending Against Model Stealing Attacks With Adaptive MisinformationSanjay Kariyappa, Moinuddin K. QureshiCVPR 2020
- Prediction Poisoning: Towards Defenses Against DNN Model Stealing AttacksTribhuvanesh Orekondy, Bernt Schiele, Mario FritzICLR 2020 · 被引用 194 次
- Adversarial Defense via Learning to Generate Diverse AttacksYunseok Jang, Tianchen Zhao, Seunghoon Hong, Honglak LeeICCV 2019 · 被引用 88 次
- On Breaking Deep Generative Model-based Defenses and BeyondYanzhi Chen, Renjie Xie, Zhanxing ZhuICML 2020 · 被引用 4 次
