Statistically Undetectable Backdoors in Deep Neural Networks
Andrej Bogdanov, Alon Rosen, Neekon Vafa
摘要
We show how an adversarial model trainer can plant backdoors in a large class of deep, feedforward neural networks. These backdoors are statistically undetectable in the white-box setting, meaning that the backdoored and honestly trained models are close in total variation distance, even given the full descriptions of the models (e.g., all of the weights). The backdoor provides access to invariance-based adversarial examples for every input, mapping distant inputs to unusually close outputs. However, without the backdoor, it is provably impossible (under standard cryptographic assumptions) to generate any such adversarial examples in polynomial time. Our theoretical and preliminary empirical findings demonstrate a fundamental power asymmetry between model trainers and model users.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Indistinguishability obfuscation from well-founded assumptionsAayush Jain, Huijia Lin, Amit SahaiSTOC 2021 · 被引用 223 次
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
- Indistinguishability Obfuscation from LPN over , DLIN, and PRGs in NC0Aayush Jain, Huijia Lin, Amit SahaiEUROCRYPT 2022 · 被引用 102 次
- Planting Undetectable Backdoors in Machine Learning Models : [Extended Abstract]Shafi Goldwasser, Michael P. Kim, Vinod Vaikuntanathan, Or ZamirFOCS 2022 · 被引用 40 次
- Frozen 1-RSB structure of the symmetric Ising perceptronWill Perkins, Changji XuSTOC 2021 · 被引用 31 次
相关 Paper
- Handcrafted Backdoors in Deep Neural NetworksSanghyun Hong, Nicholas Carlini, Alexey KurakinNeurIPS 2022 · 被引用 105 次
- Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language ModelsAlkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Katerina Sotiraki 等NeurIPS 2024 · 被引用 5 次
- Architectural Backdoors in Neural NetworksMikel Bober-Irizar, Ilia Shumailov, Yiren Zhao, Robert Mullins 等CVPR 2023
- On the Trade-off between Adversarial and Backdoor RobustnessCheng-Hsin Weng, Yan-Ting Lee, Shan-Hung WuNeurIPS 2020 · 被引用 70 次
- Black-box Detection of Backdoor Attacks with Limited Information and DataYinpeng Dong, Xiao Yang, Zhijie Deng, Tianyu Pang 等ICCV 2021 · 被引用 128 次
