OTI: A Model-free and Visually Interpretable Measure of Image Attackability
Jiaming Liang, Haowei Liu, Chi-Man Pun
摘要
Despite the tremendous success of neural networks, benign images can be corrupted by adversarial perturbations to deceive these models. Intriguingly, images differ in their attackability. Specifically, given an attack configuration, some images are easily corrupted, whereas others are more resistant. Evaluating image attackability has important applications in active learning, adversarial training, and attack enhancement. This prompts a growing interest in developing attackability measures. However, existing methods are scarce and suffer from two major limitations: (1) They rely on a model proxy to provide prior knowledge (e.g., gradients or minimal perturbation) to extract model-dependent image features. Unfortunately, in practice, many task-specific models are not readily accessible. (2) Extracted features characterizing image attackability lack visual interpretability, obscuring their direct relationship with the images. To address these, we propose a novel Object Texture Intensity (OTI), a model-free and visually interpretable measure of image attackability, which measures image attackability as the texture intensity of the image's semantic object. Theoretically, we describe the principles of OTI from the perspectives of decision boundaries as well as the mid- and high-frequency characteristics of adversarial perturbations. Comprehensive experiments demonstrate that OTI is effective and computationally efficient. In addition, our OTI provides the adversarial machine learning community with a visual understanding of attackability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao 等ICCV 2021 · 被引用 473 次
- Structure Invariant Transformation for better Adversarial TransferabilityXiaosen Wang, Zeliang Zhang, Jianping ZhangICCV 2023 · 被引用 130 次
- Revisiting Adversarial Training for ImageNet: Architectures, Training and Generalization across Threat ModelsNaman Deep Singh, Francesco Croce, Matthias HeinNeurIPS 2023 · 被引用 119 次
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang 等ICLR 2024 · 被引用 112 次
- An Adaptive Model Ensemble Adversarial Attack for Boosting Adversarial TransferabilityBin Chen, Jia-Li Yin, Shukai Chen, Bohao Chen 等ICCV 2023 · 被引用 96 次
相关 Paper
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Unrestricted Adversarial Examples via Semantic ManipulationAnand Bhattad, Min Jin Chong, Kaizhao Liang, Bo Li 等ICLR 2020 · 被引用 177 次
- How do humans perceive adversarial text? A reality check on the validity and naturalness of word-based adversarial attacksSalijona Dyrmishi, Salah Ghamizi, Maxime CordyACL 2023 · 被引用 4 次
- Robust Feature-Level Adversaries are Interpretability ToolsStephen Casper, Max Nadeau, Dylan Hadfield-Menell, Gabriel KreimanNeurIPS 2022 · 被引用 34 次
- Towards Global Explanations of Convolutional Neural Networks With Concept AttributionWeibin Wu, Yuxin Su, Xixian Chen, Shenglin Zhao 等CVPR 2020
