Self-Supervised Implicit Glyph Attention for Text Recognition
Tongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang, Qi Feng, Yudi Zhao, Wei Shen
摘要
The attention mechanism has become the de facto module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e. , implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious characterlevel bounding box annotations and would be memoryintensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SIGA). SIGA delineates the glyph structures of text images by jointly self-supervised text segmentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Self-supervised Character-to-Character Distillation for Text RecognitionTongkun Guan, Wei Shen, Xue Yang, Qi Feng 等ICCV 2023 · 被引用 36 次
- SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text RecognitionYongkun Du, Zhineng Chen, Hongtao Xie, Caiyan Jia 等ICCV 2025 · 被引用 22 次
- Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachZiyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang 等AAAI 2024 · 被引用 20 次
- Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and EditingBoqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin WangCVPR 2024 · 被引用 9 次
- CodePercept: Code-Grounded Visual STEM Perception for MLLMsTongkun Guan, Zhibo Yang, Jianqiang Wan, Mingkun Yang 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Decoupled Attention Network for Text RecognitionTianwei Wang, Yuanzhi Zhu, Lianwen Jin, Canjie Luo 等AAAI 2020 · 被引用 289 次
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang 等ICCV 2021 · 被引用 184 次
- TextScanner: Reading Characters in Order for Robust Scene Text RecognitionZhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai 等AAAI 2020 · 被引用 158 次
- PIMNet: A Parallel, Iterative and Mimicking Network for Scene Text RecognitionZhi Qiao, Yu Zhou, Jin Wei, Wei Wang 等ACM MM 2021 · 被引用 81 次
相关 Paper
- Exploring Font-independent Features for Scene Text RecognitionYizhi Wang, Zhouhui LianACM MM 2020 · 被引用 19 次
- SCATTER: Selective Context Attentional Scene Text RecognizerRon Litman, Oron Anschel, Shahar Tsiper, Roee Litman 等CVPR 2020
- Trust Prophet or Not? Taking a Further Verification Step toward Accurate Scene Text RecognitionAnna Zhu, Ke Xiao, Bo Zhou, Runmin WangACM MM 2024 · 被引用 6 次
- You Can even Annotate Text with Voice: Transcription-only-Supervised Text SpottingJingqun Tang, Su Qiao, Benlei Cui, Yuhang Ma 等ACM MM 2022 · 被引用 22 次
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun 等AAAI 2022 · 被引用 70 次
