Tracing the Origin of Adversarial Attack for Forensic Investigation and Deterrence
Han Fang, Jiyi Zhang, Yupeng Qiu, Jiayang Liu, Ke Xu, Chengfang Fang, Ee-Chien Chang
Abstract
Deep neural networks are vulnerable to adversarial attacks. In this paper, we take the role of investigators who want to trace the attack and identify the source, that is, the particular model which the adversarial examples are generated from. Techniques derived would aid forensic investigation of attack incidents and serve as deterrence to potential attacks. We consider the buyers-seller setting where a machine learning model is to be distributed to various buyers and each buyer receives a slightly different copy with the same functionality. A malicious buyer generates adversarial examples from a particular copy and uses them to attack other copies. From these adversarial examples, the investigator wants to identify the source . To address this problem, we propose a two-stage separate-and-trace framework. The model separation stage generates multiple copies of a model for the same classification task. This process injects unique features into each copy so that adversarial examples generated have distinct and traceable features. We give a parallel structure which pairs a unique tracer with the original classification model in each copy and a variational autoencoder (VAE)-based training method to achieve this goal. The tracing stage takes in adversarial examples and a few candidate models, and identifies the likely source. Based on the unique features induced by the tracer, we could effectively trace the potential adversarial copy by considering the output logits from each tracer. Empirical results show that it is possible to trace the origin of the adversarial example and the mechanism can be applied to a wide range of architectures and datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37e1c6bc-898f-479e-9e5f-4e182ff06756Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- QEBA: Query-Efficient Boundary-Based Blackbox AttackHuichen Li, Xiaojun Xu, Xiaolu Zhang, Shuang Yang et al.CVPR 2020
- SurFree: A Fast Surrogate-Free Black-Box AttackThibault Maho, Teddy Furon, Erwan Le MerrerCVPR 2021
Related papers
- Identification of the Adversary from a Single Adversarial ExampleMinhao Cheng, Rui Min, Haochen Sun, Pin-Yu ChenICML 2023 · 1 citation
- SepMark: Deep Separable Watermarking for Unified Source Tracing and Deepfake DetectionXiaoshuai Wu, Xin Liao, Bo OuACM MM 2023 · 74 citations
- DeepTracer: Tracing Stolen Model via Deep Coupled WatermarksYunfei Yang, Xiaojun Chen, Yuexin Xuan, Zhendong Zhao et al.AAAI 2026
- False Claims against Model Ownership ResolutionJian Liu, Rui Zhang, Sebastian Szyller, Kui Ren et al.USENIX Security 2024 · 22 citations
- Deep Neural Network Fingerprinting by Conferrable Adversarial ExamplesNils Lukas, Yuxuan Zhang, Florian KerschbaumICLR 2021 · 182 citations
