Beyond Pretrained Features: Noisy Image Modeling Provides Adversarial Defense
Zunzhi You, Daochang Liu, Bohyung Han, Chang Xu
摘要
Recent advancements in masked image modeling (MIM) have made it a prevailing framework for self-supervised visual representation learning. The MIM pretrained models, like most deep neural network methods, remain vulnerable to adversarial attacks, limiting their practical application, and this issue has received little research attention. In this paper, we investigate how this powerful self-supervised learning paradigm can provide adversarial robustness to downstream classifiers. During the exploration, we find that noisy image modeling (NIM), a simple variant of MIM that adopts denoising as the pre-text task, reconstructs noisy images surprisingly well despite severe corruption. Motivated by this observation, we propose an adversarial defense method, referred to as De 3 , by exploiting the pretrained decoder for denoising. Through De 3 , NIM is able to enhance adversarial robustness beyond providing pretrained features. Furthermore, we incorporate a simple modification, sampling the noise scale hyperparameter from random distributions, and enable the defense to achieve a better and tunable trade-off between accuracy and robustness. Experimental results demonstrate that, in terms of adversarial robustness, NIM is superior to MIM thanks to its effective denoising capability. Moreover, the defense provided by NIM achieves performance on par with adversarial training while offering the extra tunability advantage. Source code and models are available at https://github.com/youzunzhi/NIM-AdvDef .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MIMIR: Masked Image Modeling for Mutual Information-based Adversarial RobustnessXiaoyun Xu, Shujian Yu, Zhuoran Liu, Stjepan PicekNDSS 2026 · 被引用 12 次
- Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual RepresentationWenzhao Xiang, Chang Liu, Hongyang Yu, Xilin ChenAAAI 2025 · 被引用 3 次
- Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language ModelsJia-Wei Hai, Yijun Wang, Xiu-Shen WeiICML 2026
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Removing Adversarial Noise in Class Activation Feature SpaceDawei Zhou, Nannan Wang, Chunlei Peng, Xinbo Gao 等ICCV 2021 · 被引用 37 次
- Ditch the Denoiser: Emergence of Noise Robustness in Self-Supervised Learning from Data CurriculumWenquan Lu, Jiaqi Zhang, Hugues Van Assel, Randall BalestrieroNeurIPS 2025 · 被引用 5 次
- Adversarial Masking for Self-Supervised LearningYuge Shi, N. Siddharth, Philip H. S. Torr, Adam R. KosiorekICML 2022 · 被引用 110 次
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 被引用 15 次
- DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative DenoisingZhenhao Li, Huichi Zhou, Marek Rei, Lucia SpeciaACL 2025
