Cracking White-box DNN Watermarks via Invariant Neuron Transforms
Xudong Pan, Mi Zhang, Yifan Yan, Yining Wang, Min Yang
摘要
Recently, how to protect the Intellectual Property (IP) of deep neural networks (DNN) becomes a major concern for the AI industry. To combat potential model piracy, recent works explore various watermarking strategies to embed secret identity messages into the prediction behaviors or the internals (e.g., weights and neuron activation) of the target model. Sacrificing less functionality and involving more knowledge about the target model, the latter branch of watermarking schemes (i.e., whitebox model watermarking) is claimed to be accurate, credible and secure against most known watermark removal attacks, with emerging research efforts and applications in the industry. In this paper, we present the first effective removal attack which cracks almost all the existing white-box watermarking schemes with provably no performance overhead and no required prior knowledge. By analyzing these IP protection mechanisms at the granularity of neurons, we for the first time discover their common dependence on a set of fragile features of a local neuron group, all of which can be arbitrarily tampered by our proposed chain of invariant neuron transforms. On 9 state-of-the-art whitebox watermarking schemes and a broad set of industry-level DNN architectures, our attack for the first time reduces the embedded identity message in the protected models to be almost random. Meanwhile, unlike known removal attacks, our attack requires no prior knowledge on the training data distribution or the adopted watermark algorithms, and leaves model functionality intact.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesTianlong Xu, Chen Wang, Gaoyang Liu, Yang Yang 等NeurIPS 2024 · 被引用 17 次
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li 等ACM MM 2024 · 被引用 14 次
- Dimension-independent Certified Neural Network Watermarks via Mollifier SmoothingJiaxiang Ren, Yang Zhou, Jiayin Jin, Lingjuan Lyu 等ICML 2023 · 被引用 10 次
它引用的顶会 Paper20
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang 等NDSS 2019 · 被引用 1,141 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
相关 Paper
- Rethinking White-Box Watermarks on Deep Learning Models under Neural Structural ObfuscationYifan Yan, Xudong Pan, Mi Zhang, Min YangUSENIX Security 2023
- Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited DataYifan Lu, Wenxuan Li, Mi Zhang, Xudong Pan 等CCS 2024 · 被引用 2 次
- DeepEclipse: How to Break White-Box DNN-Watermarking SchemesAlessandro Pegoraro, Carlotta Segna, Kavita Kumari, Ahmad-Reza SadeghiUSENIX Security 2024 · 被引用 11 次
- Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?Run Wang, Haoxuan Li, Lingzhou Mu, Jixing Ren 等ACM MM 2022 · 被引用 9 次
- IPRemover: A Generative Model Inversion Attack against Deep Neural Network Fingerprinting and WatermarkingWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek 等AAAI 2024 · 被引用 12 次
