MSA2: Multi-Task Framework With Structure-Aware and Style-Adaptive Character Representation for Open-Set Chinese Text Recognition
Yangfu Li, Hongjian Zhan, Qi Liu, Li Sun, Yu-Jie Xiong, Yue Lu
摘要
Most existing methods regard open-set Chinese text recognition (CTR) as a single-task problem, primarily focusing on prototype learning of linguistic components or glyphs to identify unseen characters. In contrast, humans identify characters by integrating multiple perspectives, including linguistic and visual cues. Inspired by this, we propose a multi-task framework termed MSA 2 , which considers multi-view character representations for open-set CTR. Within MSA 2 , we introduce two novel strategies for character representation: structure-aware component encoding (SACE) and style-adaptive glyph embedding (SAGE). SACE utilizes a binary tree with dynamic representation space to emphasize the primary linguistic components, thereby generating structure-aware and discriminative linguistic representations for each character. Meanwhile, SAGE employs glyph-centric contrastive learning to aggregate features from diverse forms, yielding robust glyph representations for the CTR model to adapt to the style variations among various fonts. Extensive experiments demonstrate that our proposed MSA 2 outperforms state-of-the-art CTR methods, achieving average improvements of 1.3% and 6.0% in accuracy under closed-set and open-set settings on the BCTR dataset, respectively. The code is available at https://github.com/LPAIS/MSA-2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chinese Text Recognition with A Pre-Trained CLIP-Like Model Through Image-IDS AligningHaiyang Yu, Xiaocong Wang, Bin Li, Xiangyang XueICCV 2023 · 被引用 43 次
- Open-Set Text Recognition via Character-Context DecouplingChang Liu, Chun Yang, Xu-Cheng YinCVPR 2022 · 被引用 33 次
- Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text RecognitionShancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao 等CVPR 2021
- SEED: Semantics Enhanced Encoder-Decoder Framework for Scene Text RecognitionZhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou 等CVPR 2020
相关 Paper
- Context-Based Contrastive Learning for Scene Text RecognitionXinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun 等AAAI 2022 · 被引用 70 次
- Few-shot Font Generation with Localized Style Representations and FactorizationSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等AAAI 2021 · 被引用 111 次
- XMP-Font: Self-Supervised Cross-Modality Pre-training for Few-Shot Font GenerationWei Liu, Fangyue Liu, Fei Ding, Qian He 等CVPR 2022 · 被引用 64 次
- Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text RecognitionHao Liu, Bin Wang, Zhimin Bao, Mobai Xue 等AAAI 2022 · 被引用 49 次
- Multiple Heads are Better than One: Few-shot Font Generation with Multiple Localized ExpertsSong Park, Sanghyuk Chun, Junbum Cha, Bado Lee 等ICCV 2021 · 被引用 96 次
