Exploring Effective Data for Surrogate Training Towards Black-box Attack
Xuxiang Sun, Gong Cheng, Hongda Li, Lei Pei, Junwei Han
Abstract
Without access to the training data where a black-box victim model is deployed, training a surrogate model for black-box adversarial attack is still a struggle. In terms of data, we mainly identify three key measures for effective surrogate training in this paper. First, we show that leveraging the loss introduced in this paper to enlarge the inter-class similarity makes more sense than enlarging the inter-class diversity like existing methods. Next, unlike the approaches that expand the intra-class diversity in an implicit model-agnostic fashion, we propose a loss function specific to the surrogate model for our generator to enhance the intra-class diversity. Finally, in accordance with the in-depth observations for the methods based on proxy data, we argue that leveraging the proxy data is still an effective way for surrogate training. To this end, we propose a triple-player framework by introducing a discriminator into the traditional data-free framework. In this way, our method can be competitive when there are few semantic overlaps between the scarce proxy data (with the size between 1 k and 5k) and the training data. We evaluate our method on a range of victim models and datasets. The extensive results witness the effectiveness of our method. Our source code is available at https://github.com/xuxiangsun/ST-Data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50281a82-b8d1-4722-b695-c19f1fbb8aeeCited by top-tier papers3
- Exploring Query Efficient Data Generation Towards Data-Free Model Stealing in Hard Label SettingGaozheng Pei, Shaojie Lyu, Ke Ma, Pinci Yang et al.AAAI 2025 · 2 citations
- Adversarial Robustness via Random Projection FiltersMinjing Dong, Chang XuCVPR 2023
- PGA: Prior-free Generative Attack for Practical No-box Scenariohongyu peng, Xiang Yuan, Gong ChengCVPR 2026
Builds on21
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNetsDongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey et al.ICLR 2020 · 357 citations
Related papers
- Delving into Data: Effectively Substitute Training for Black-box AttackWenxuan Wang, Bangjie Yin, Taiping Yao, Li Zhang et al.CVPR 2021
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 31 citations
- DaST: Data-Free Substitute Training for Adversarial AttacksMingyi Zhou, Jing Wu, Yipeng Liu, Shuaicheng Liu et al.CVPR 2020
- PFFAA: Prototype-based Feature and Frequency Alteration Attack for Semantic SegmentationZhidong Yu, Zhenbo Shi, Xiaoman Liu, Wei YangACM MM 2024 · 2 citations
- Dual Student Networks for Data-Free Model StealingJames Beetham, Navid Kardan, Ajmal Saeed Mian, Mubarak ShahICLR 2023 · 3 citations
