What Machines See Is Not What They Get: Fooling Scene Text Recognition Models With Adversarial Text Images
Xing Xu, Jiefu Chen, Jinhui Xiao, Lianli Gao, Fumin Shen, Heng Tao Shen
摘要
The research on scene text recognition (STR) has made remarkable progress in recent years with the development of deep neural networks (DNNs). Recent studies on adversarial attack have verified that a DNN model designed for non-sequential tasks (e.g., classification, segmentation and retrieval) can be easily fooled by adversarial examples. Actually, STR is an application highly related to security issues. However, there are few studies considering the safety and reliability of STR models that make sequential prediction. In this paper, we make the first attempt in attacking the state-of-the-art DNN-based STR models. Specifically, we propose a novel and efficient optimization-based method that can be naturally integrated to different sequential prediction schemes, i.e., connectionist temporal classification (CTC) and attention mechanism. We apply our proposed method to five state-of-the-art STR models with both targeted and untargeted attack modes, the comprehensive results on 7 real-world datasets and 2 synthetic datasets consistently show the vulnerability of these STR models with a significant performance drop. Finally, we also test our attack method on a real-world STR engine of Baidu OCR, which demonstrates the practical potentials of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 被引用 33 次
- Towards the Unseen: Iterative Text Recognition by Distilling from ErrorsAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yi-Zhe SongICCV 2021 · 被引用 19 次
- An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareWenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen 等ASE 2023 · 被引用 7 次
- Text's Armor: Optimized Local Adversarial Perturbation Against Scene Text Editing AttacksTao Xiang, Hangcheng Liu, Shangwei Guo, Hantao Liu 等ACM MM 2022 · 被引用 5 次
- Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character RecognitionJiacheng Deng, Li Dong, Jiahao Chen, Diqun Yan 等ACM MM 2023 · 被引用 5 次
它引用的顶会 Paper5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park 等ICCV 2019 · 被引用 551 次
- Towards Adversarially Robust Object DetectionHaichao Zhang, Jianyu WangICCV 2019 · 被引用 152 次
- Stealthy Adversarial Perturbations Against Real-Time Video Classification SystemsShasha Li, Ajaya Neupane, Sujoy Paul, Chengyu Song 等NDSS 2019 · 被引用 132 次
- Universal Perturbation Attack Against Image RetrievalJie Li, Rongrong Ji, Hong Liu, Xiaopeng Hong 等ICCV 2019 · 被引用 115 次
相关 Paper
- Learning Optimization-based Adversarial Perturbations for Attacking Sequential Recognition ModelsXing Xu, Jiefu Chen, Jinhui Xiao, Zheng Wang 等ACM MM 2020 · 被引用 18 次
- Towards Irreversible Attack: Fooling Scene Text Recognition via Multi-Population Coevolution SearchJingyu Li, Pengwen Dai, Mingqing Zhu, Chengwei Wang 等NeurIPS 2025
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi 等AAAI 2020 · 被引用 151 次
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia 等ICCV 2025 · 被引用 22 次
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai 等AAAI 2020 · 被引用 158 次
