What Machines See Is Not What They Get: Fooling Scene Text Recognition Models With Adversarial Text Images
Xing Xu, Jiefu Chen, Jinhui Xiao, Lianli Gao, Fumin Shen, Heng Tao Shen
Abstract
The research on scene text recognition (STR) has made remarkable progress in recent years with the development of deep neural networks (DNNs). Recent studies on adversarial attack have verified that a DNN model designed for non-sequential tasks (e.g., classification, segmentation and retrieval) can be easily fooled by adversarial examples. Actually, STR is an application highly related to security issues. However, there are few studies considering the safety and reliability of STR models that make sequential prediction. In this paper, we make the first attempt in attacking the state-of-the-art DNN-based STR models. Specifically, we propose a novel and efficient optimization-based method that can be naturally integrated to different sequential prediction schemes, i.e., connectionist temporal classification (CTC) and attention mechanism. We apply our proposed method to five state-of-the-art STR models with both targeted and untargeted attack modes, the comprehensive results on 7 real-world datasets and 2 synthetic datasets consistently show the vulnerability of these STR models with a significant performance drop. Finally, we also test our attack method on a real-world STR engine of Baidu OCR, which demonstrates the practical potentials of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge DistillationAyan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury, Yi-Zhe SongICCV 2021 · 33 citations
- Towards the Unseen: Iterative Text Recognition by Distilling from ErrorsAyan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yi-Zhe SongICCV 2021 · 19 citations
- An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareWenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen et al.ASE 2023 · 7 citations
- Text's Armor: Optimized Local Adversarial Perturbation Against Scene Text Editing AttacksTao Xiang, Hangcheng Liu, Shangwei Guo, Hantao Liu et al.ACM MM 2022 · 5 citations
- Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character RecognitionJiacheng Deng, Li Dong, Jiahao Chen, Diqun Yan et al.ACM MM 2023 · 5 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park et al.ICCV 2019 · 551 citations
- Towards Adversarially Robust Object DetectionHaichao Zhang, Jianyu WangICCV 2019 · 152 citations
- Stealthy Adversarial Perturbations Against Real-Time Video Classification SystemsShasha Li, Ajaya Neupane, Sujoy Paul, Chengyu Song et al.NDSS 2019 · 132 citations
- Universal Perturbation Attack Against Image RetrievalJie Li, Rongrong Ji, Hong Liu, Xiaopeng Hong et al.ICCV 2019 · 115 citations
Related papers
- Learning Optimization-based Adversarial Perturbations for Attacking Sequential Recognition ModelsXing Xu, Jiefu Chen, Jinhui Xiao, Zheng Wang et al.ACM MM 2020 · 18 citations
- Towards Irreversible Attack: Fooling Scene Text Recognition via Multi-Population Coevolution SearchJingyu Li, Pengwen Dai, Mingqing Zhu, Chengwei Wang et al.NeurIPS 2025
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia et al.ICCV 2025 · 22 citations
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai et al.AAAI 2020 · 158 citations
