Identification of the Adversary from a Single Adversarial Example
Minhao Cheng, Rui Min, Haochen Sun, Pin-Yu Chen
Abstract
Deep neural networks have been shown vulnerable to adversarial examples. Even though many defense methods have been proposed to enhance the robustness, it is still a long way toward providing an attack-free method to build a trustworthy machine learning system. In this paper, instead of enhancing the robustness, we take the investigator's perspective and propose a new framework to trace the first compromised model copy in a forensic investigation manner. Specifically, we focus on the following setting: the machine learning service provider provides model copies for a set of customers. However, one of the customers conducted adversarial attacks to fool the system. Therefore, the investigator's objective is to identify the first compromised copy by collecting and analyzing evidence from only available adversarial examples. To make the tracing viable, we design a random mask watermarking mechanism to differentiate adversarial examples from different copies. First, we propose a tracing approach in the data-limited case where the original example is also available. Then, we design a data-free approach to identify the adversary without accessing the original example. Finally, the effectiveness of our proposed framework is evaluated by extensive experiments with different model architectures, adversarial attacks, and datasets. Our code is publicly available at https://github.com/ rmin2000/adv_tracing.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Sign-OPT: A Query-Efficient Hard-label Adversarial AttackMinhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen et al.ICLR 2020 · 256 citations
Related papers
- Tracing the Origin of Adversarial Attack for Forensic Investigation and DeterrenceHan Fang, Jiyi Zhang, Yupeng Qiu, Jiayang Liu et al.ICCV 2023 · 3 citations
- DeepTracer: Tracing Stolen Model via Deep Coupled WatermarksYunfei Yang, Xiaojun Chen, Yuexin Xuan, Zhendong Zhao et al.AAAI 2026
- RIGA: Covert and Robust White-Box Watermarking of Deep Neural NetworksTianhao Wang, Florian KerschbaumWWW 2021 · 128 citations
- False Claims against Model Ownership ResolutionJian Liu, Rui Zhang, Sebastian Szyller, Kui Ren et al.USENIX Security 2024 · 22 citations
- SepMark: Deep Separable Watermarking for Unified Source Tracing and Deepfake DetectionXiaoshuai Wu, Xin Liao, Bo OuACM MM 2023 · 74 citations
