MEA-Defender: A Robust Watermark against Model Extraction Attack
Peizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou, Shengzhi Zhang, Ruigang Liang, Shenchen Zhu, Pan Li, Yingjun Zhang
摘要
Recently, numerous highly-valuable Deep Neural Networks (DNNs) have been trained using deep learning algorithms. To protect the Intellectual Property (IP) of the original owners over such DNN models, backdoor-based watermarks have been extensively studied. However, most of such watermarks fail upon model extraction attack, which utilizes input samples to query the target model and obtains the corresponding outputs, thus training a substitute model using such input-output pairs. In this paper, we propose a novel watermark to protect IP of DNN models against model extraction, named MEA-Defender. In particular, we obtain the watermark by combining two samples from two source classes in the input domain and design a watermark loss function that makes the output domain of the watermark within that of the main task samples. Since both the input domain and the output domain of our watermark are indispensable parts of those of the main task samples, the watermark will be extracted into the stolen model along with the main task during model extraction. We conduct extensive experiments on four model extraction attacks, using five datasets and six models trained based on supervised learning and self-supervised learning algorithms. The experimental results demonstrate that MEA-Defender is highly robust against different model extraction attacks, and various watermark removal/detection approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li 等ACM MM 2024 · 被引用 14 次
- HoneypotNet: Backdoor Attacks Against Model ExtractionYixu Wang, Tianle Gu, Yan Teng, Yingchun Wang 等AAAI 2025 · 被引用 4 次
- "Yes, My LoRD." Guiding Language Model Extraction with Locality Reinforced DistillationZi Liang, Qingqing Ye, Yanyun Wang, Sen Zhang 等ACL 2025 · 被引用 2 次
- Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model DistributionZhicheng Zhang, Peizhuo Lv, Mengke Wan, Jiang Fang 等ACM MM 2025 · 被引用 2 次
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai 等NDSS 2026 · 被引用 1 次
它引用的顶会 Paper20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
相关 Paper
- Deep Neural Network Watermarking against Model Extraction AttackJingxuan Tan, Nan Zhong, Zhenxing Qian, Xinpeng Zhang 等ACM MM 2023 · 被引用 34 次
- IPRemover: A Generative Model Inversion Attack against Deep Neural Network Fingerprinting and WatermarkingWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek 等AAAI 2024 · 被引用 12 次
- Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?Run Wang, Haoxuan Li, Lingzhou Mu, Jixing Ren 等ACM MM 2022 · 被引用 9 次
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 被引用 18 次
- Identification for Deep Neural Network: Simply Adjusting Few Weights!Yingjie Lao, Peng Yang, Weijie Zhao, Ping LiICDE 2022 · 被引用 19 次
