CAPatch: Physical Adversarial Patch against Image Captioning Systems
Shibo Zhang, Yushi Cheng, Wenjun Zhu, Xiaoyu Ji, Wenyuan Xu
摘要
The fast-growing surveillance systems will make image captioning, i.e., automatically generating text descriptions of images, an essential technique to process the huge volumes of videos efficiently, and correct captioning is essential to ensure the text authenticity. While prior work has demonstrated the feasibility of fooling computer vision models with adversarial patches, it is unclear whether the vulnerability can lead to incorrect captioning, which involves natural language processing after image feature extraction. In this paper, we design CAPatch, a physical adversarial patch that can result in mistakes in the final captions, i.e., either create a completely different sentence or a sentence with keywords missing, against multi-modal image captioning systems. To make CAPatch effective and practical in the physical world, we propose a detection assurance and attention enhancement method to increase the impact of CAPatch and a robustness improvement method to address the patch distortions caused by image printing and capturing. Evaluations on three commonly-used image captioning systems (Show-and-Tell, Self-critical Sequence Training: Att2in, and Bottom-up Top-down) demonstrate the effectiveness of CAPatch in both the digital and physical worlds, whereby volunteers wear printed patches in various scenarios, clothes, lighting conditions. With a size of 5% of the image, physically-printed CAPatch can achieve continuous attacks with an attack success rate higher than 73.1% over a video recorder.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person DetectionHui Wei, Zhixiang Wang, Kewei Zhang, Jiaqi Hou 等NeurIPS 2024 · 被引用 22 次
- AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement LearningVasudev Gohil, Satwik Patnaik, Dileep Kalathil, Jeyavijayan RajendranUSENIX Security 2024 · 被引用 9 次
- PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency ReasoningZhiyuan Yu, Ao Li, Ruoyao Wen, Yijia Chen 等CCS 2024 · 被引用 4 次
- MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM AgentsYun Xing, Nhat Chung, Jie Zhang, Yue Cao 等AAAI 2026
- Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsSong Yan, Hui Wei, Jinlong Fei, Guoliang Yang 等ACM MM 2025
它引用的顶会 Paper9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 被引用 1,295 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
相关 Paper
- TPatch: A Triggered Physical Adversarial PatchWenjun Zhu, Xiaoyu Ji, Yushi Cheng, Shibo Zhang 等USENIX Security 2023
- Legitimate Adversarial Patches: Evading Human Eyes and Detection Models in the Physical WorldJia Tan, Nan Ji, Haidong Xie, Xueshuang XiangACM MM 2021 · 被引用 44 次
- Physically Adversarial Infrared Patches with Learnable Shapes and LocationsXingxing Wei, Jie Yu, Yao HuangCVPR 2023
- Adversarial Texture for Fooling Person Detectors in the Physical WorldZhanhao Hu, Siyuan Huang, Xiaopei Zhu, Fuchun Sun 等CVPR 2022 · 被引用 125 次
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image ClassifierChong Xiang, Saeed Mahloujifar, Prateek MittalUSENIX Security 2022
