Consensus-Robust Transfer Attacks via Parameter and Representation Perturbations
Shixin Li, Zewei Li, Xiaojing Ma, Xiaofan Bai, Pingyi Hu, Dongmei Zhang, Bin Zhu
摘要
Adversarial examples crafted on one model often exhibit poor transferability to others, hindering their effectiveness in black-box settings. This limitation arises from two key factors: (i) decision-boundary variation across models and (ii) representation drift in feature space. We address these challenges through a new perspective that frames transferability for untargeted attacks as a consensus-robust optimization problem: adversarial perturbations should remain effective across a neighborhood of plausible target models. To model this uncertainty, we introduce two complementary perturbation channels: a parameter channel , capturing boundary shifts via weight perturbations, and a representation channel , addressing feature drift via stochastic blending of clean and adversarial activations. We then propose CORTA (COnsensus–Robust Transfer Attack), a lightweight attack instantiated from this robust formulation using two first-order strategies: (i) sensitivity regularization based on the squared Frobenius norm of logits’ Jacobian with respect to weights, and (ii) Monte Carlo sampling for blended feature representations. Our theoretical analysis provides a certified lower bound linking these approximations to the robust objective. Extensive experiments on CIFAR-100 and ImageNet show that CORTA significantly outperforms state-of-the-art transfer-based methods— including ensemble approaches—across CNN and Vision Transformer targets. Notably, CORTA achieves a 19.1 percentage-point gain in transfer success rate over the best prior method while using only a single surrogate model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang 等ICLR 2020 · 被引用 765 次
相关 Paper
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang 等ICLR 2024 · 被引用 112 次
- DRIFT: Divergent Response in Filtered Transformations for Robust Adversarial DefenseAmira Guesmi, Muhammad ShafiqueICLR 2026 · 被引用 1 次
- Improving Black-Box Generative Attacks via Generator Semantic ConsistencyJongoh Jeong, Hunmin Yang, Jaeseok Jeong, Kuk-Jin YoonICLR 2026
- Training Meta-Surrogate Model for Transferable Adversarial AttackYunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Cho-Jui HsiehAAAI 2023 · 被引用 31 次
- RaPA: Enhancing Transferable Targeted Attacks via Random Parameter PruningTongrui Su, Qingbin Li, Shengyu Zhu, Wei Chen 等CVPR 2026 · 被引用 1 次
