Knowledge Mining with Scene Text for Fine-Grained Recognition
Hao Wang, Junchao Liao, Tianheng Cheng, Zewen Gao, Hao Liu, Bo Ren, Xiang Bai, Wenyu Liu
Abstract
Recently, the semantics of scene text has been proven to be essential in fine-grained image classification. However, the existing methods mainly exploit the literal meaning of scene text for fine-grained recognition, which might be irrelevant when it is not significantly related to objects/scenes. We propose an end-to-end trainable network that mines implicit contextual knowledge behind scene text image and enhance the semantics and correlation to fine-tune the image representation. Unlike the existing methods, our model integrates three modalities: visual feature extraction, text semantics extraction, and correlating background knowledge to fine-grained image classification. Specifically, we employ KnowBert to retrieve relevant knowledge for semantic representation and combine it with image features for fine-grained classification. Experiments on two benchmark datasets, Con-Text, and Drink Bottle, show that our method outperforms the state-of-the-art by 3.72% mAP and 5.39% mAp, respectively. To further validate the effectiveness of the proposed method, we create a new dataset on crowd activity recognition for the evaluation. The source code and new dataset of this work are available at this repository <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> https://github.com/lanfeng4659/KnowledgeMiningWithSceneText.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Towards Robust Real-Time Scene Text Detection: From Semantic to Instance Representation LearningXugong Qin, Pengyuan Lyu, Chengquan Zhang, Yu Zhou et al.ACM MM 2023 · 21 citations
- Self-Distillation Regularized Connectionist Temporal Classification Loss for Text Recognition: A Simple Yet Effective ApproachZiyin Zhang, Ning Lu, Minghui Liao, Yongshuai Huang et al.AAAI 2024 · 20 citations
- Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language ModelsJeonghwan Kim, Heng JiEMNLP 2024 · 4 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
- All You Need Is Boundary: Toward Arbitrary-Shaped Text SpottingHao Wang, Pu Lu, Hui Zhang, Mingkun Yang et al.AAAI 2020 · 145 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
- Scene Text Retrieval via Joint Text Detection and Similarity LearningHao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu et al.CVPR 2021
Related papers
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionShancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao et al.CVPR 2021
- Towards Accurate Scene Text Recognition With Semantic Reasoning NetworksDeli Yu, Xuan Li, Chengquan Zhang, Tao Liu et al.CVPR 2020
- Vision-Language Pre-Training for Boosting Scene Text DetectorsSibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang et al.CVPR 2022 · 38 citations
- Self-Attention Based Text Knowledge Mining for Text DetectionQi Wan, Haoqin Ji, Linlin ShenCVPR 2021
- GTC: Guided Training of CTC towards Efficient and Accurate Scene Text RecognitionWenyang Hu, Xiaocong Cai, Jun Hou, Shuai Yi et al.AAAI 2020 · 151 citations
