USENIX Security2023Top-tier venue
CAPatch: Physical Adversarial Patch against Image Captioning Systems
Shibo Zhang, Yushi Cheng, Wenjun Zhu, Xiaoyu Ji, Wenyuan Xu
Abstract
The fast-growing surveillance systems will make image captioning, i.e., automatically generating text descriptions of images, an essential technique to process the huge volumes of videos efficiently, and correct captioning is essential to ensure the text authenticity. While prior work has demonstrated the feasibility of fooling computer vision models with adversarial patches, it is unclear whether the vulnerability can lead to incorrect captioning, which involves natural language processing after image feature extraction. In this paper, we design CAPatch, a physical adversarial patch that can result in mistakes in the final captions, i.e., either create a completely different sentence or a sentence with keywords missing, against multi-modal image captioning systems. To make CAPatch effective and practical in the physical world, we propose a detection assurance and attention enhancement method to increase the impact of CAPatch and a robustness improvement method to address the patch distortions caused by image printing and capturing. Evaluations on three commonly-used image captioning systems (Show-and-Tell, Self-critical Sequence Training: Att2in, and Bottom-up Top-down) demonstrate the effectiveness of CAPatch in both the digital and physical worlds, whereby volunteers wear printed patches in various scenarios, clothes, lighting conditions. With a size of 5% of the image, physically-printed CAPatch can achieve continuous attacks with an attack success rate higher than 73.1% over a video recorder.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8111660-692d-4f87-a732-9850bed9fe95Cited by top-tier papers6
- Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person DetectionHui Wei, Zhixiang Wang, Kewei Zhang, Jiaqi Hou et al.NeurIPS 2024 · 22 citations
- AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement LearningVasudev Gohil, Satwik Patnaik, Dileep Kalathil, Jeyavijayan RajendranUSENIX Security 2024 · 9 citations
- PhySense: Defending Physically Realizable Attacks for Autonomous Systems via Consistency ReasoningZhiyuan Yu, Ao Li, Ruoyao Wen, Yijia Chen et al.CCS 2024 · 4 citations
- MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM AgentsYun Xing, Nhat Chung, Jie Zhang, Yue Cao et al.AAAI 2026
- Universally Unfiltered and Unseen: Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsSong Yan, Hui Wei, Jinlong Fei, Guoliang Yang et al.ACM MM 2025
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
Related papers
- TPatch: A Triggered Physical Adversarial PatchWenjun Zhu, Xiaoyu Ji, Yushi Cheng, Shibo Zhang et al.USENIX Security 2023
- Legitimate Adversarial Patches: Evading Human Eyes and Detection Models in the Physical WorldJia Tan, Nan Ji, Haidong Xie, Xueshuang XiangACM MM 2021 · 44 citations
- Physically Adversarial Infrared Patches with Learnable Shapes and LocationsXingxing Wei, Jie Yu, Yao HuangCVPR 2023
- Adversarial Texture for Fooling Person Detectors in the Physical WorldZhanhao Hu, Siyuan Huang, Xiaopei Zhu, Fuchun Sun et al.CVPR 2022 · 125 citations
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image ClassifierChong Xiang, Saeed Mahloujifar, Prateek MittalUSENIX Security 2022
