Training Meta-Surrogate Model for Transferable Adversarial Attack
Yunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui Hsieh
Abstract
We consider adversarial attacks to a black-box model when no queries are allowed. In this setting, many methods directly attack surrogate models and transfer the obtained adversarial examples to fool the target model. Plenty of previous works investigated what kind of attacks to the surrogate model can generate more transferable adversarial examples, but their performances are still limited due to the mismatches between surrogate models and the target model. In this paper, we tackle this problem from a novel angleinstead of using the original surrogate models, can we obtain a Meta-Surrogate Model (MSM) such that attacks to this model can be easier transferred to other models? We show that this goal can be mathematically formulated as a well-posed (bi-level-like) optimization problem and design a differentiable attacker to make training feasible. Given one or a set of surrogate models, our method can thus obtain an MSM such that adversarial examples generated on MSM enjoy eximious transferability. Comprehensive experiments on Cifar-10 and ImageNet demonstrate that by attacking the MSM, we can obtain stronger transferable adversarial examples to fool black-box models including adversarially trained ones, with much higher success rates than existing methods. The proposed method reveals significant security challenges of deep models and is promising to be served as a state-of-the-art benchmark for evaluating the robustness of deep models in the black-box setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Non-Adaptive Adversarial Face GenerationSunpill Kim, Seunghun Paik, Chanwoo Hwang, Minsu Kim et al.NeurIPS 2025 · 5 citations
- Taxonomy Driven Fast Adversarial TrainingKun Tong, Chengze Jiang, Jie Gui, Yuan CaoAAAI 2024 · 2 citations
- Boosting Adversarial Transferability via Ensemble Non-AttentionYipeng Zou, Qin Liu, Jie Wu, Yu Peng et al.AAAI 2026
Builds on16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
Related papers
- LRS: Enhancing Adversarial Transferability through Lipschitz Regularized SurrogateTao Wu, Tie Luo, Donald C. Wunsch IIAAAI 2024 · 11 citations
- Boosting Black-Box Attack with Partially Transferred Conditional Adversarial DistributionYan Feng, Baoyuan Wu, Yanbo Fan, Li Liu et al.CVPR 2022 · 34 citations
- Towards Multiple Black-boxes Attack via Adversarial Example Generation NetworkMingxing Duan, Kenli Li, Lingxi Xie, Qi Tian et al.ACM MM 2021 · 21 citations
- Minimizing Maximum Model Discrepancy for Transferable Black-box Targeted AttacksAnqi Zhao, Tong Chu, Yahao Liu, Wen Li et al.CVPR 2023
- Meta Gradient Adversarial AttackZheng Yuan, Jie Zhang, Yunpei Jia, Chuanqi Tan et al.ICCV 2021 · 95 citations
