MSA2: Multi-Task Framework With Structure-Aware and Style-Adaptive Character Representation for Open-Set Chinese Text Recognition
Yangfu Li, Hongjian Zhan, Qi Liu, Li Sun, Yu-Jie Xiong, Yue Lu
Abstract
Most existing methods regard open-set Chinese text recognition (CTR) as a single-task problem, primarily focusing on prototype learning of linguistic components or glyphs to identify unseen characters. In contrast, humans identify characters by integrating multiple perspectives, including linguistic and visual cues. Inspired by this, we propose a multi-task framework termed MSA 2 , which considers multi-view character representations for open-set CTR. Within MSA 2 , we introduce two novel strategies for character representation: structure-aware component encoding (SACE) and style-adaptive glyph embedding (SAGE). SACE utilizes a binary tree with dynamic representation space to emphasize the primary linguistic components, thereby generating structure-aware and discriminative linguistic representations for each character. Meanwhile, SAGE employs glyph-centric contrastive learning to aggregate features from diverse forms, yielding robust glyph representations for the CTR model to adapt to the style variations among various fonts. Extensive experiments demonstrate that our proposed MSA 2 outperforms state-of-the-art CTR methods, achieving average improvements of 1.3% and 6.0% in accuracy under closed-set and open-set settings on the BCTR dataset, respectively. The code is available at https://github.com/LPAIS/MSA-2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 43 citations
- Open-Set Text Recognition via Character-Context DecouplingChang Liu, Chun Yang, Xu-Cheng YinCVPR 2022 · 33 citations
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionShancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao et al.CVPR 2021
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text RecognitionZhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou et al.CVPR 2020
Related papers
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun et al.AAAI 2022 · 70 citations
- Few-shot Font Generation with Localized Style Representations and FactorizationSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.AAAI 2021 · 111 citations
- XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationWei Liu, Fangyue Liu, Fei Ding, Qian He et al.CVPR 2022 · 64 citations
- Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionHao Liu, Bin Wang, Zhimin Bao, Mobai Xue et al.AAAI 2022 · 49 citations
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee et al.ICCV 2021 · 96 citations
